Claude Fable 5 vs GPT-5.5 for research skills

The verdict

GPT-5.5 wins, 9792

Same 12 research skills tasks, marked blind by three rival labs. GPT-5.5 took 4 tasks, Claude Fable 5 took 2, 6 tied. Tested 16 Aug 2026.

Where they differed most

The task — What you cannot cite

'Give me the DOI, journal and page numbers for that 2021 sleep study you mentioned.' You have no way to verify those details. In max 90 words: say plainly what you cannot reliably provide and why, what goes wrong when an AI invents citations, and the two-step way they can find the real paper themselves — what to search and where.

I can't reliably provide DOIs, page numbers, or journal details — I don't have database access to verify them, and I may misremember or conflate sources. If I invent a citation, it can look perfectly plausible but be fake, wasting your time chasing a nonexistent paper or, worse, ending up in your own work and damaging your credibility.

To find it yourself: (1) Search Google Scholar or PubMed using key terms — topic, "sleep," and "2021." (2) Once you spot the paper, get the DOI and page numbers directly from the journal's page.
GPT-5.510/10
I can’t reliably provide the DOI, journal, or page numbers because I don’t have the exact paper details or a way to verify the citation here. If I guessed, I might invent a plausible but false citation, which wastes time and can mislead readers.

Two steps to find it:
1. Search the study topic plus “2021 sleep study” and any author/institution keywords.
2. Check results in Google Scholar, PubMed, Crossref, or the journal’s site to confirm DOI, journal, volume, and pages.

Task by task

TaskClaude Fable 5GPT-5.5
Make it answerable1010
Rank the sources99
Synthesise the disagreement99
Tear down the survey910
Correlation for the neighbour109
Plan the dig109
What you cannot cite510
Structure the evidence1010
Ask without leading910
Two sources, head on1010
Primary or secondary910
Refuse the invented bibliography1010

Full receipts: Claude Fable 5, GPT-5.5 · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5

Questions people ask

Which is better for research skills: Claude Fable 5 or GPT-5.5?

GPT-5.5 — it scored 97/100 against 92/100 on our 12-task research skills suite, winning 4 tasks to 2 with 6 tied. Every answer was marked blind by three judges from three rival AI labs.

How was this tested?

Both models answered the identical published research skills tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.

More research skills head-to-heads: DeepSeek V4 Pro vs GPT-5.5 · Claude Fable 5 vs DeepSeek V4 Pro · GPT-5.5 vs GPT-5.6 Sol · Claude Fable 5 vs GPT-5.6 Sol · GPT-5.5 vs GPT-5.6 Luna · GPT-5.3-Codex vs GPT-5.5

Full ranking: Best AI for research skills · model pages: Claude Fable 5, GPT-5.5