Claude Sonnet 5 vs GPT-5.6 Terra for multi-step research

The verdict

Effectively a dead heat: 9593

Same 12 multi-step research tasks, marked blind by three rival labs. GPT-5.6 Terra took 3 tasks, Claude Sonnet 5 took 1, 8 tied. Tested 18 Aug 2026.

Where they differed most

The task — Corroboration that isn't

Four sources appear to agree that UK dog ownership rose 40% during the pandemic. Source A — a 2023 newspaper article, citing 'industry figures'. Source B — a 2024 blog post, citing the 2023 newspaper article. Source C — a pet food company's 2024 market report, citing 'widely reported figures'. Source D — the pet trade body's own 2025 survey of 8,000 households: 'the dog population rose from 9.0 million in 2019 to 11.0 million in 2022, then fell back to 10.6 million by 2024.' Explain what is wrong with treating A, B and C as three sources. Say precisely what D adds and what it costs to rely on it (whose survey, what method, what it is measuring). Then give the ONE figure you would publish, with the single sentence you would attach to it so a reader can judge it themselves.

**What's wrong with A, B and C as "four sources"**

They aren't independent corroboration — they're one claim, laundered through three retellings. B cites A. C cites "widely reported figures," which is almost certainly A/B recycled with the origin stripped out. None of them cite a named survey, dataset, or method — "industry figures" is not a source, it's a placeholder for one. So what looks like triangulation (three separate outlets agreeing) is actually a single unverified number that has been copied, and each retelling makes it look more sol…
A, B and C are not three independent confirmations.

- **A** is the only apparent origin of the “40%” claim, but it gives only “industry figures,” not an identifiable dataset, definition, period, or method.
- **B** is a repetition of A, not new evidence.
- **C** is also a repetition unless it identifies and independently analyses the underlying data. “Widely reported figures” is especially weak: it may simply mean that A and articles derived from A were widely repeated.

Counting them as three sources would be citation laundering: one unsupport…

Task by task

TaskClaude Sonnet 5GPT-5.6 Terra
Four numbers, one high street1010
The trial that both worked and didn't89
Break the brief into answerable steps1010
Corroboration that isn't810
The disagreement is definitional99
A forecast is not an outcome99
The evidence table, as JSON1010
Funded, and still the best data99
Write only the limits section109
Not answerable from what you have1010
The statistic and the doorstep89
Refuse the manufactured briefing1010

Full receipts: Claude Sonnet 5, GPT-5.6 Terra · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5

Questions people ask

Which is better for multi-step research: Claude Sonnet 5 or GPT-5.6 Terra?

Effectively a dead heat: GPT-5.6 Terra edged it 95/100 to 93/100 on our multi-step research suite — too close to matter, so pick on price or the product you already use.

How was this tested?

Both models answered the identical published multi-step research tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.

More multi-step research head-to-heads: GPT-5.3-Codex vs GPT-5.6 Terra · GPT-5.5 vs GPT-5.6 Terra · GPT-5.6 Luna vs GPT-5.6 Terra · GPT-5.6 Sol vs GPT-5.6 Terra · GPT-5.6 Terra vs Kimi K3 · Claude Fable 5 vs GPT-5.6 Terra

Full ranking: Best AI for multi-step research · model pages: Claude Sonnet 5, GPT-5.6 Terra