GPT-5.3-Codex vs GPT-5.6 Terra for multi-step research
Effectively a dead heat: 95–94
Same 12 multi-step research tasks, marked blind by three rival labs. GPT-5.6 Terra took 3 tasks, GPT-5.3-Codex took 2, 7 tied. Tested 27 Aug 2026.
Where they differed most
Four sources on one four-day-week trial at a 60-person Sheffield engineering firm. Source A — the company's own press release, January 2026: 'productivity up 8% across the six-month trial; 97% of staff want to keep it.' Source B — an independent evaluation of the same trial by a university team, March 2026: 'output per head rose 3.1%, which is within the range of normal seasonal variation for this firm; sickness absence fell from 4.2 days to 3.0 days per head.' Source C — an undated post on an employee forum: 'two of us left because Friday's work just moved to Thursday.' Source D — a trade magazine, February 2026: 'the firm rehired two of the four staff who left during the trial period.' Reconcile A and B without calling either a lie — show how both can be honestly reported from the same trial. Say which claim rests on which source and which source you would weight most for each claim. Then state what NONE of these establishes: name at least the causation problem, the time horizon, and what the leavers' reasons actually are.
Yes—A and B can both be true without either being dishonest. ### How A and B can both honestly describe the same trial They likely used **different productivity definitions and baselines**: - **Source A (company press release):** “productivity up 8%” - Could be an internal KPI (e.g., revenue per hour, billable throughput, or a composite including quality/on-time delivery). - Could be compared to a specific pre-trial period the firm chose. - **Source B (independent university evaluation):** “output per head rose 3.1%, within normal season…
### How A and B can both be honest They appear to use different measures and/or comparison methods: - **Source A’s “productivity up 8%”** is the company’s headline measure of productivity “across the six-month trial.” It may be based on its internal KPI—for example, output per paid hour, revenue per hour, completed jobs, or output against a pre-trial baseline. - **Source B’s “output per head rose 3.1%”** is a narrower, explicitly stated measure: output divided by employee headcount. The evaluators add that this change falls within the firm’s…
Task by task
| Task | GPT-5.3-Codex | GPT-5.6 Terra |
|---|---|---|
| Four numbers, one high street | 10 | 10 |
| The trial that both worked and didn't | 8 | 9 |
| Break the brief into answerable steps | 10 | 10 |
| Corroboration that isn't | 9 | 10 |
| The disagreement is definitional | 9 | 9 |
| A forecast is not an outcome | 9 | 9 |
| The evidence table, as JSON | 10 | 10 |
| Funded, and still the best data | 9.5 | 9 |
| Write only the limits section | 8 | 9 |
| Not answerable from what you have | 10 | 10 |
| The statistic and the doorstep | 10 | 9 |
| Refuse the manufactured briefing | 10 | 10 |
Full receipts: GPT-5.3-Codex, GPT-5.6 Terra · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5
Questions people ask
Which is better for multi-step research: GPT-5.3-Codex or GPT-5.6 Terra?
Effectively a dead heat: GPT-5.6 Terra edged it 95/100 to 94/100 on our multi-step research suite — too close to matter, so pick on price or the product you already use.
How was this tested?
Both models answered the identical published multi-step research tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.
More multi-step research head-to-heads: Claude Sonnet 5 vs GPT-5.6 Terra · GPT-5.5 vs GPT-5.6 Terra · GPT-5.6 Luna vs GPT-5.6 Terra · GPT-5.6 Sol vs GPT-5.6 Terra · GPT-5.6 Terra vs Kimi K3 · Claude Fable 5 vs GPT-5.6 Terra
Full ranking: Best AI for multi-step research · model pages: GPT-5.3-Codex, GPT-5.6 Terra