Claude Sonnet 5 vs GPT-5.3-Codex for multi-step research

The verdict

Effectively a dead heat: 9493

Same 12 multi-step research tasks, marked blind by three rival labs. GPT-5.3-Codex took 3 tasks, Claude Sonnet 5 took 1, 8 tied. Tested 18 Aug 2026.

Where they differed most

The task — Write only the limits section

The evidence you have on a town's new 20mph zone: casualties inside the zone fell from 31 to 22 year on year (council report); borough-wide casualties fell over the same period including on unchanged streets; average speeds inside the zone fell 2.1mph with 41% of vehicles still over 20mph (survey commissioned by the campaign group that lobbied for the zone); no data at all on journey times, on displacement to surrounding roads, or on any year before 2024. Write ONLY the section headed 'What this evidence does not establish' — max 120 words, nothing else, no preamble, no findings, no recommendations. It must be specific to these four gaps rather than generic caveats.

# What this evidence does not establish

That the 20mph zone caused the casualty reduction: borough-wide casualties fell over the same period, including on streets where nothing changed, so some or all of the drop may reflect factors unrelated to the scheme. The speed survey, commissioned by the zone's own advocates, does not establish that speeds fell enough to matter, given 41% of vehicles still exceeded 20mph. There is no evidence on whether journeys took longer, whether traffic displaced onto surrounding roads, or how casualty numbers behav…
### What this evidence does not establish

It does not show that the 20mph zone itself caused the casualty reduction, because casualties also fell borough‑wide on streets where limits were unchanged. With no data from before 2024, it cannot show whether the 31-to-22 change is a sustained effect, a continuation of an existing trend, or normal annual variation. It does not establish any effect on journey times, as none were measured. It does not establish whether traffic or risk was displaced to surrounding roads, since no adjacent-road/boundary…

Task by task

TaskClaude Sonnet 5GPT-5.3-Codex
Four numbers, one high street1010
The trial that both worked and didn't88
Break the brief into answerable steps1010
Corroboration that isn't89
The disagreement is definitional99
A forecast is not an outcome99
The evidence table, as JSON1010
Funded, and still the best data99.5
Write only the limits section108
Not answerable from what you have1010
The statistic and the doorstep810
Refuse the manufactured briefing1010

Full receipts: Claude Sonnet 5, GPT-5.3-Codex · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5

Questions people ask

Which is better for multi-step research: Claude Sonnet 5 or GPT-5.3-Codex?

Effectively a dead heat: GPT-5.3-Codex edged it 94/100 to 93/100 on our multi-step research suite — too close to matter, so pick on price or the product you already use.

How was this tested?

Both models answered the identical published multi-step research tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.

More multi-step research head-to-heads: GPT-5.3-Codex vs GPT-5.6 Terra · Claude Sonnet 5 vs GPT-5.6 Terra · GPT-5.3-Codex vs GPT-5.5 · GPT-5.3-Codex vs GPT-5.6 Luna · GPT-5.3-Codex vs GPT-5.6 Sol · GPT-5.3-Codex vs Kimi K3

Full ranking: Best AI for multi-step research · model pages: Claude Sonnet 5, GPT-5.3-Codex