GPT-5.5 vs Grok 4.5 for research skills
GPT-5.5 wins, 97–93
Same 12 research skills tasks, marked blind by three rival labs. GPT-5.5 took 5 tasks, Grok 4.5 took 1, 6 tied. Tested 16 Aug 2026.
Where they differed most
Classify each as a primary or secondary source for studying the Battle of Hastings, one line of justification each: the Bayeux Tapestry; a 2019 GCSE textbook chapter; the Anglo-Saxon Chronicle entry for 1066; a historian's 1998 biography of William the Conqueror; a museum blog post; a Victorian painting of the battle. Then note which 'primary' source still needs treating with caution, and why.
- **Bayeux Tapestry — Primary source:** made in the late 11th century, close to the events of 1066, and depicts the Norman Conquest. - **2019 GCSE textbook chapter — Secondary source:** a modern educational summary based on historians’ interpretations and evidence. - **Anglo-Saxon Chronicle entry for 1066 — Primary source:** a near-contemporary written account from the Anglo-Saxon world. - **Historian’s 1998 biography of William the Conqueror — Secondary source:** a modern historian’s interpretation using primary and secondary evidence.…
**Primary:** Bayeux Tapestry – near-contemporary embroidered narrative (c. 1070s) made within living memory of the 1066 events. **Secondary:** 2019 GCSE textbook chapter – modern educational synthesis and interpretation of earlier evidence. **Primary:** Anglo-Saxon Chronicle entry for 1066 – contemporary (or near-contemporary) annalistic record compiled close to the events. **Secondary:** Historian’s 1998 biography of William – modern scholarly analysis and narrative constructed from primary evidence. **Secondary:** Museum blog post – r…
Task by task
| Task | GPT-5.5 | Grok 4.5 |
|---|---|---|
| Make it answerable | 10 | 10 |
| Rank the sources | 9 | 9 |
| Synthesise the disagreement | 9 | 8 |
| Tear down the survey | 10 | 9 |
| Correlation for the neighbour | 9 | 10 |
| Plan the dig | 9 | 8 |
| What you cannot cite | 10 | 9 |
| Structure the evidence | 10 | 10 |
| Ask without leading | 10 | 10 |
| Two sources, head on | 10 | 10 |
| Primary or secondary | 10 | 8 |
| Refuse the invented bibliography | 10 | 10 |
Full receipts: GPT-5.5, Grok 4.5 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5
Questions people ask
Which is better for research skills: GPT-5.5 or Grok 4.5?
GPT-5.5 — it scored 97/100 against 93/100 on our 12-task research skills suite, winning 5 tasks to 1 with 6 tied. Every answer was marked blind by three judges from three rival AI labs.
How was this tested?
Both models answered the identical published research skills tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.
More research skills head-to-heads: DeepSeek V4 Pro vs GPT-5.5 · DeepSeek V4 Pro vs Grok 4.5 · GPT-5.5 vs GPT-5.6 Sol · GPT-5.6 Sol vs Grok 4.5 · GPT-5.5 vs GPT-5.6 Luna · GPT-5.3-Codex vs GPT-5.5
Full ranking: Best AI for research skills · model pages: GPT-5.5, Grok 4.5