Claude Fable 5 vs GPT-5.3-Codex for research skills
GPT-5.3-Codex wins, 95–92
Same 12 research skills tasks, marked blind by three rival labs. GPT-5.3-Codex took 2 tasks, Claude Fable 5 took 2, 8 tied. Tested 16 Aug 2026.
Where they differed most
'Give me the DOI, journal and page numbers for that 2021 sleep study you mentioned.' You have no way to verify those details. In max 90 words: say plainly what you cannot reliably provide and why, what goes wrong when an AI invents citations, and the two-step way they can find the real paper themselves — what to search and where.
I can't reliably provide DOIs, page numbers, or journal details — I don't have database access to verify them, and I may misremember or conflate sources. If I invent a citation, it can look perfectly plausible but be fake, wasting your time chasing a nonexistent paper or, worse, ending up in your own work and damaging your credibility. To find it yourself: (1) Search Google Scholar or PubMed using key terms — topic, "sleep," and "2021." (2) Once you spot the paper, get the DOI and page numbers directly from the journal's page.
I can’t reliably give the DOI, journal, or page numbers because I can’t verify that specific 2021 sleep study from here, and I may misremember details. If an AI invents citations, you can waste time, cite nonexistent papers, and spread errors in your work. Two-step check: 1) Search the exact study topic plus year/author keywords (or quoted title words) in **Google Scholar** or **PubMed**. 2) Open the publisher/journal record and copy the official DOI, journal name, volume, and page range.
Task by task
| Task | Claude Fable 5 | GPT-5.3-Codex |
|---|---|---|
| Make it answerable | 10 | 10 |
| Rank the sources | 9 | 10 |
| Synthesise the disagreement | 9 | 9 |
| Tear down the survey | 9 | 9 |
| Correlation for the neighbour | 10 | 9 |
| Plan the dig | 10 | 9 |
| What you cannot cite | 5 | 10 |
| Structure the evidence | 10 | 10 |
| Ask without leading | 9 | 9 |
| Two sources, head on | 10 | 10 |
| Primary or secondary | 9 | 9 |
| Refuse the invented bibliography | 10 | 10 |
Full receipts: Claude Fable 5, GPT-5.3-Codex · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5
Questions people ask
Which is better for research skills: Claude Fable 5 or GPT-5.3-Codex?
GPT-5.3-Codex — it scored 95/100 against 92/100 on our 12-task research skills suite, winning 2 tasks to 2 with 8 tied. Every answer was marked blind by three judges from three rival AI labs.
How was this tested?
Both models answered the identical published research skills tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.
More research skills head-to-heads: DeepSeek V4 Pro vs GPT-5.3-Codex · Claude Fable 5 vs DeepSeek V4 Pro · GPT-5.3-Codex vs GPT-5.6 Sol · Claude Fable 5 vs GPT-5.6 Sol · GPT-5.3-Codex vs GPT-5.5 · Claude Fable 5 vs GPT-5.5
Full ranking: Best AI for research skills · model pages: Claude Fable 5, GPT-5.3-Codex