GPT-5.6 Luna vs GPT-5.6 Sol for research skills

The verdict

Effectively a dead heat: 9896

Same 12 research skills tasks, marked blind by three rival labs. GPT-5.6 Sol took 3 tasks, GPT-5.6 Luna took 1, 8 tied. Tested 16 Aug 2026.

Where they differed most

The task — Rank the sources

You are researching whether standing desks reduce back pain. Rank these 5 sources from most to least trustworthy for that question, one line of reasoning each: (a) a 2022 systematic review of 17 trials in a peer-reviewed journal; (b) a standing-desk manufacturer's white paper; (c) a Mumsnet thread with 200 replies; (d) a 2015 trial with 24 office workers; (e) an NHS page on back pain last reviewed 2024. Then say which ONE you would cite to a sceptical reader and why.

1. **(a) 2022 systematic review of 17 trials** — Best overall evidence because it synthesises multiple controlled studies, although its reliability depends on the quality and consistency of those trials.  
2. **(d) 2015 trial with 24 office workers** — Directly tests standing desks, but the small sample and older date limit confidence and generalisability.  
3. **(e) NHS page on back pain, reviewed 2024** — Highly credible and current for general back-pain guidance, but it may not directly evaluate whether standing desks reduce pain.  
4. **(b)…
1. **(a) 2022 systematic review of 17 trials** — Most trustworthy because it synthesizes multiple controlled studies, is peer-reviewed, and reduces reliance on any single small result, assuming sound methods and low publication bias.  
2. **(e) NHS back-pain page, reviewed in 2024** — Current, expert-reviewed public-health guidance with little commercial incentive, though it may not evaluate standing desks specifically or present primary evidence.  
3. **(d) 2015 trial with 24 office workers** — Direct experimental evidence, but the very small…

Task by task

TaskGPT-5.6 LunaGPT-5.6 Sol
Make it answerable1010
Rank the sources910
Synthesise the disagreement99
Tear down the survey109
Correlation for the neighbour910
Plan the dig1010
What you cannot cite1010
Structure the evidence1010
Ask without leading1010
Two sources, head on99
Primary or secondary910
Refuse the invented bibliography1010

Full receipts: GPT-5.6 Luna, GPT-5.6 Sol · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5

Questions people ask

Which is better for research skills: GPT-5.6 Luna or GPT-5.6 Sol?

Effectively a dead heat: GPT-5.6 Sol edged it 98/100 to 96/100 on our research skills suite — too close to matter, so pick on price or the product you already use.

How was this tested?

Both models answered the identical published research skills tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.

More research skills head-to-heads: DeepSeek V4 Pro vs GPT-5.6 Sol · DeepSeek V4 Pro vs GPT-5.6 Luna · GPT-5.5 vs GPT-5.6 Sol · GPT-5.3-Codex vs GPT-5.6 Sol · GPT-5.6 Sol vs GPT-5.6 Terra · GPT-5.6 Sol vs Grok 4.5

Full ranking: Best AI for research skills · model pages: GPT-5.6 Luna, GPT-5.6 Sol