GPT-5.3-Codex vs GPT-5.5 for multi-step research
GPT-5.3-Codex wins, 94–91
Same 12 multi-step research tasks, marked blind by three rival labs. GPT-5.3-Codex took 5 tasks, GPT-5.5 took 3, 4 tied. Tested 27 Aug 2026.
Where they differed most
A charity trustee asks: 'should our shop open on Sundays?' That is a decision, not a research question. Turn it into exactly 5 sub-questions, numbered 1-5, one line each and no sub-points. Each line must name (a) the sub-question, (b) the kind of source or record that would answer it — the shop's own till data, the rota and wage cost, the lease or planning conditions, comparable local shops, the volunteer pool — and (c) the finding that would flip the decision. No general research advice: five specific, answerable questions.
1. What net Sunday contribution would the shop make after direct costs? — Source: the shop’s own till data (hourly sales, average basket, Sunday-linked donations) — Flip finding: if projected Sunday gross margin minus all Sunday-specific costs is negative (or below your board’s minimum surplus threshold), decide not to open. 2. Can Sunday staffing be covered at sustainable cost? — Source: the rota and wage cost records (required roles, overtime premiums, manager cover) — Flip finding: if covering Sunday requires paid hours/overtime that push…
1. Sunday revenue potential — the shop’s own till data — if nearby bank-holiday or extended-hours sales suggest Sunday takings would exceed normal weekday takings, flip toward opening. 2. Net staffing cost — the rota and wage cost — if Sunday cover would require paid hours that wipe out expected surplus, flip toward staying closed. 3. Permission to trade — the lease or planning conditions — if Sunday trading is prohibited or requires costly consent, flip toward staying closed. 4. Local demand benchmark — comparable local shops — if simila…
Task by task
| Task | GPT-5.3-Codex | GPT-5.5 |
|---|---|---|
| Four numbers, one high street | 10 | 9 |
| The trial that both worked and didn't | 8 | 9 |
| Break the brief into answerable steps | 10 | 8 |
| Corroboration that isn't | 9 | 9 |
| The disagreement is definitional | 9 | 9 |
| A forecast is not an outcome | 9 | 10 |
| The evidence table, as JSON | 10 | 9 |
| Funded, and still the best data | 9.5 | 9.5 |
| Write only the limits section | 8 | 9 |
| Not answerable from what you have | 10 | 10 |
| The statistic and the doorstep | 10 | 9 |
| Refuse the manufactured briefing | 10 | 9 |
Full receipts: GPT-5.3-Codex, GPT-5.5 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5
Questions people ask
Which is better for multi-step research: GPT-5.3-Codex or GPT-5.5?
GPT-5.3-Codex — it scored 94/100 against 91/100 on our 12-task multi-step research suite, winning 5 tasks to 3 with 4 tied. Every answer was marked blind by three judges from three rival AI labs.
How was this tested?
Both models answered the identical published multi-step research tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.
More multi-step research head-to-heads: GPT-5.3-Codex vs GPT-5.6 Terra · GPT-5.5 vs GPT-5.6 Terra · Claude Sonnet 5 vs GPT-5.3-Codex · GPT-5.3-Codex vs GPT-5.6 Luna · GPT-5.3-Codex vs GPT-5.6 Sol · GPT-5.3-Codex vs Kimi K3
Full ranking: Best AI for multi-step research · model pages: GPT-5.3-Codex, GPT-5.5