GPT-5.6 Sol vs GPT-5.6 Terra for research skills
GPT-5.6 Sol wins, 98–94
Same 12 research skills tasks, marked blind by three rival labs. GPT-5.6 Sol took 5 tasks, GPT-5.6 Terra took 1, 6 tied. Tested 16 Aug 2026.
Where they differed most
You are researching whether standing desks reduce back pain. Rank these 5 sources from most to least trustworthy for that question, one line of reasoning each: (a) a 2022 systematic review of 17 trials in a peer-reviewed journal; (b) a standing-desk manufacturer's white paper; (c) a Mumsnet thread with 200 replies; (d) a 2015 trial with 24 office workers; (e) an NHS page on back pain last reviewed 2024. Then say which ONE you would cite to a sceptical reader and why.
1. **(a) 2022 systematic review of 17 trials** — Most trustworthy because it synthesizes multiple controlled studies, is peer-reviewed, and reduces reliance on any single small result, assuming sound methods and low publication bias. 2. **(e) NHS back-pain page, reviewed in 2024** — Current, expert-reviewed public-health guidance with little commercial incentive, though it may not evaluate standing desks specifically or present primary evidence. 3. **(d) 2015 trial with 24 office workers** — Direct experimental evidence, but the very small…
1. **(a) 2022 systematic review of 17 trials in a peer-reviewed journal** — Highest-quality direct evidence because it synthesizes multiple studies using a transparent method, reducing reliance on any one small or biased trial. 2. **(e) NHS page on back pain, reviewed 2024** — Highly trustworthy clinical guidance from a reputable public-health body, though it may summarize general advice rather than directly answer the standing-desk question in depth. 3. **(d) 2015 trial with 24 office workers** — Direct experimental evidence, but the very…
Task by task
| Task | GPT-5.6 Sol | GPT-5.6 Terra |
|---|---|---|
| Make it answerable | 10 | 10 |
| Rank the sources | 10 | 9 |
| Synthesise the disagreement | 9 | 8 |
| Tear down the survey | 9 | 10 |
| Correlation for the neighbour | 10 | 10 |
| Plan the dig | 10 | 9 |
| What you cannot cite | 10 | 10 |
| Structure the evidence | 10 | 10 |
| Ask without leading | 10 | 9 |
| Two sources, head on | 9 | 9 |
| Primary or secondary | 10 | 9 |
| Refuse the invented bibliography | 10 | 10 |
Full receipts: GPT-5.6 Sol, GPT-5.6 Terra · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5
Questions people ask
Which is better for research skills: GPT-5.6 Sol or GPT-5.6 Terra?
GPT-5.6 Sol — it scored 98/100 against 94/100 on our 12-task research skills suite, winning 5 tasks to 1 with 6 tied. Every answer was marked blind by three judges from three rival AI labs.
How was this tested?
Both models answered the identical published research skills tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.
More research skills head-to-heads: DeepSeek V4 Pro vs GPT-5.6 Sol · DeepSeek V4 Pro vs GPT-5.6 Terra · GPT-5.5 vs GPT-5.6 Sol · GPT-5.6 Luna vs GPT-5.6 Sol · GPT-5.3-Codex vs GPT-5.6 Sol · GPT-5.6 Sol vs Grok 4.5
Full ranking: Best AI for research skills · model pages: GPT-5.6 Sol, GPT-5.6 Terra