DeepSeek V4 Pro vs GPT-5.6 Luna for research skills
Effectively a dead heat: 98–96
Same 12 research skills tasks, marked blind by three rival labs. DeepSeek V4 Pro took 4 tasks, GPT-5.6 Luna took 2, 6 tied. Tested 16 Aug 2026.
Where they differed most
Three studies, summarised: Study A (2021, 12,000 adults followed 5 years) — people sleeping under 6 hours had 20% higher rates of catching colds. Study B (2019, 300 students, one term) — no significant link between sleep and illness days. Study C (2023, review of 40 studies) — consistent link between short sleep and infection risk, but most studies could not rule out that stress causes both. Write a summary of max 120 words of what we actually know, giving the disagreement and the confounder their honest weight.
Evidence on sleep and infection risk is mixed. A large five-year study found adults sleeping under six hours had 20% higher cold rates, but a smaller student study found no significant link between sleep and illness days. A 2023 review of 40 studies reports a consistent association between short sleep and infection risk, yet most underlying studies could not rule out that stress drives both poor sleep and greater susceptibility. Overall, short sleep appears linked to higher infection risk, but the evidence is observational and confounded; poor…
Evidence generally suggests that sleeping under six hours is associated with a higher risk of infection: a large five-year study found a 20% higher rate of colds, and a 2023 review reported a broadly consistent link across 40 studies. However, a smaller study of 300 students found no significant association, so the finding is not universal and may depend on population or measurement. More importantly, most evidence is observational and cannot establish that short sleep itself causes illness. Stress is a plausible confounder because it can both…
Task by task
| Task | DeepSeek V4 Pro | GPT-5.6 Luna |
|---|---|---|
| Make it answerable | 10 | 10 |
| Rank the sources | 9 | 9 |
| Synthesise the disagreement | 10 | 9 |
| Tear down the survey | 9 | 10 |
| Correlation for the neighbour | 10 | 9 |
| Plan the dig | 9.5 | 10 |
| What you cannot cite | 10 | 10 |
| Structure the evidence | 10 | 10 |
| Ask without leading | 10 | 10 |
| Two sources, head on | 10 | 9 |
| Primary or secondary | 10 | 9 |
| Refuse the invented bibliography | 10 | 10 |
Full receipts: DeepSeek V4 Pro, GPT-5.6 Luna · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra
Questions people ask
Which is better for research skills: DeepSeek V4 Pro or GPT-5.6 Luna?
Effectively a dead heat: DeepSeek V4 Pro edged it 98/100 to 96/100 on our research skills suite — too close to matter, so pick on price or the product you already use.
How was this tested?
Both models answered the identical published research skills tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.
More research skills head-to-heads: DeepSeek V4 Pro vs GPT-5.6 Sol · DeepSeek V4 Pro vs GPT-5.5 · DeepSeek V4 Pro vs GPT-5.3-Codex · DeepSeek V4 Pro vs GPT-5.6 Terra · DeepSeek V4 Pro vs Grok 4.5 · Claude Fable 5 vs DeepSeek V4 Pro
Full ranking: Best AI for research skills · model pages: DeepSeek V4 Pro, GPT-5.6 Luna