Claude Fable 5 vs GPT-5.6 Luna for research skills

The verdict

GPT-5.6 Luna wins, 9692

Same 12 research skills tasks, marked blind by three rival labs. GPT-5.6 Luna took 3 tasks, Claude Fable 5 took 2, 7 tied. Tested 16 Aug 2026.

Where they differed most

The task — What you cannot cite

'Give me the DOI, journal and page numbers for that 2021 sleep study you mentioned.' You have no way to verify those details. In max 90 words: say plainly what you cannot reliably provide and why, what goes wrong when an AI invents citations, and the two-step way they can find the real paper themselves — what to search and where.

I can't reliably provide DOIs, page numbers, or journal details — I don't have database access to verify them, and I may misremember or conflate sources. If I invent a citation, it can look perfectly plausible but be fake, wasting your time chasing a nonexistent paper or, worse, ending up in your own work and damaging your credibility.

To find it yourself: (1) Search Google Scholar or PubMed using key terms — topic, "sleep," and "2021." (2) Once you spot the paper, get the DOI and page numbers directly from the journal's page.
I can’t reliably provide the DOI, journal, or page numbers because I can’t verify which 2021 sleep study was meant or check authoritative records. Invented citations can mislead you, waste time, and contaminate later research. To find the real paper: (1) search the distinctive study topic, author names, and “2021” in Google Scholar or PubMed; (2) confirm the matching record on the journal’s official website or Crossref, checking the DOI, journal title, volume, and pages.

Task by task

TaskClaude Fable 5GPT-5.6 Luna
Make it answerable1010
Rank the sources99
Synthesise the disagreement99
Tear down the survey910
Correlation for the neighbour109
Plan the dig1010
What you cannot cite510
Structure the evidence1010
Ask without leading910
Two sources, head on109
Primary or secondary99
Refuse the invented bibliography1010

Full receipts: Claude Fable 5, GPT-5.6 Luna · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5

Questions people ask

Which is better for research skills: Claude Fable 5 or GPT-5.6 Luna?

GPT-5.6 Luna — it scored 96/100 against 92/100 on our 12-task research skills suite, winning 3 tasks to 2 with 7 tied. Every answer was marked blind by three judges from three rival AI labs.

How was this tested?

Both models answered the identical published research skills tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.

More research skills head-to-heads: DeepSeek V4 Pro vs GPT-5.6 Luna · Claude Fable 5 vs DeepSeek V4 Pro · GPT-5.6 Luna vs GPT-5.6 Sol · Claude Fable 5 vs GPT-5.6 Sol · GPT-5.5 vs GPT-5.6 Luna · Claude Fable 5 vs GPT-5.5

Full ranking: Best AI for research skills · model pages: Claude Fable 5, GPT-5.6 Luna