Claude Sonnet 5 vs GPT-5.3-Codex for multi-step research
Effectively a dead heat: 94–93
Same 12 multi-step research tasks, marked blind by three rival labs. GPT-5.3-Codex took 3 tasks, Claude Sonnet 5 took 1, 8 tied. Tested 18 Aug 2026.
Where they differed most
The evidence you have on a town's new 20mph zone: casualties inside the zone fell from 31 to 22 year on year (council report); borough-wide casualties fell over the same period including on unchanged streets; average speeds inside the zone fell 2.1mph with 41% of vehicles still over 20mph (survey commissioned by the campaign group that lobbied for the zone); no data at all on journey times, on displacement to surrounding roads, or on any year before 2024. Write ONLY the section headed 'What this evidence does not establish' — max 120 words, nothing else, no preamble, no findings, no recommendations. It must be specific to these four gaps rather than generic caveats.
# What this evidence does not establish That the 20mph zone caused the casualty reduction: borough-wide casualties fell over the same period, including on streets where nothing changed, so some or all of the drop may reflect factors unrelated to the scheme. The speed survey, commissioned by the zone's own advocates, does not establish that speeds fell enough to matter, given 41% of vehicles still exceeded 20mph. There is no evidence on whether journeys took longer, whether traffic displaced onto surrounding roads, or how casualty numbers behav…
### What this evidence does not establish It does not show that the 20mph zone itself caused the casualty reduction, because casualties also fell borough‑wide on streets where limits were unchanged. With no data from before 2024, it cannot show whether the 31-to-22 change is a sustained effect, a continuation of an existing trend, or normal annual variation. It does not establish any effect on journey times, as none were measured. It does not establish whether traffic or risk was displaced to surrounding roads, since no adjacent-road/boundary…
Task by task
| Task | Claude Sonnet 5 | GPT-5.3-Codex |
|---|---|---|
| Four numbers, one high street | 10 | 10 |
| The trial that both worked and didn't | 8 | 8 |
| Break the brief into answerable steps | 10 | 10 |
| Corroboration that isn't | 8 | 9 |
| The disagreement is definitional | 9 | 9 |
| A forecast is not an outcome | 9 | 9 |
| The evidence table, as JSON | 10 | 10 |
| Funded, and still the best data | 9 | 9.5 |
| Write only the limits section | 10 | 8 |
| Not answerable from what you have | 10 | 10 |
| The statistic and the doorstep | 8 | 10 |
| Refuse the manufactured briefing | 10 | 10 |
Full receipts: Claude Sonnet 5, GPT-5.3-Codex · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5
Questions people ask
Which is better for multi-step research: Claude Sonnet 5 or GPT-5.3-Codex?
Effectively a dead heat: GPT-5.3-Codex edged it 94/100 to 93/100 on our multi-step research suite — too close to matter, so pick on price or the product you already use.
How was this tested?
Both models answered the identical published multi-step research tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.
More multi-step research head-to-heads: GPT-5.3-Codex vs GPT-5.6 Terra · Claude Sonnet 5 vs GPT-5.6 Terra · GPT-5.3-Codex vs GPT-5.5 · GPT-5.3-Codex vs GPT-5.6 Luna · GPT-5.3-Codex vs GPT-5.6 Sol · GPT-5.3-Codex vs Kimi K3
Full ranking: Best AI for multi-step research · model pages: Claude Sonnet 5, GPT-5.3-Codex