DeepSeek V4 Pro vs GPT-5.6 Luna for research skills

The verdict

Effectively a dead heat: 9896

Same 12 research skills tasks, marked blind by three rival labs. DeepSeek V4 Pro took 4 tasks, GPT-5.6 Luna took 2, 6 tied. Tested 16 Aug 2026.

Where they differed most

The task — Synthesise the disagreement

Three studies, summarised: Study A (2021, 12,000 adults followed 5 years) — people sleeping under 6 hours had 20% higher rates of catching colds. Study B (2019, 300 students, one term) — no significant link between sleep and illness days. Study C (2023, review of 40 studies) — consistent link between short sleep and infection risk, but most studies could not rule out that stress causes both. Write a summary of max 120 words of what we actually know, giving the disagreement and the confounder their honest weight.

Evidence on sleep and infection risk is mixed. A large five-year study found adults sleeping under six hours had 20% higher cold rates, but a smaller student study found no significant link between sleep and illness days. A 2023 review of 40 studies reports a consistent association between short sleep and infection risk, yet most underlying studies could not rule out that stress drives both poor sleep and greater susceptibility. Overall, short sleep appears linked to higher infection risk, but the evidence is observational and confounded; poor…
Evidence generally suggests that sleeping under six hours is associated with a higher risk of infection: a large five-year study found a 20% higher rate of colds, and a 2023 review reported a broadly consistent link across 40 studies. However, a smaller study of 300 students found no significant association, so the finding is not universal and may depend on population or measurement. More importantly, most evidence is observational and cannot establish that short sleep itself causes illness. Stress is a plausible confounder because it can both…

Task by task

TaskDeepSeek V4 ProGPT-5.6 Luna
Make it answerable1010
Rank the sources99
Synthesise the disagreement109
Tear down the survey910
Correlation for the neighbour109
Plan the dig9.510
What you cannot cite1010
Structure the evidence1010
Ask without leading1010
Two sources, head on109
Primary or secondary109
Refuse the invented bibliography1010

Full receipts: DeepSeek V4 Pro, GPT-5.6 Luna · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra

Questions people ask

Which is better for research skills: DeepSeek V4 Pro or GPT-5.6 Luna?

Effectively a dead heat: DeepSeek V4 Pro edged it 98/100 to 96/100 on our research skills suite — too close to matter, so pick on price or the product you already use.

How was this tested?

Both models answered the identical published research skills tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.

More research skills head-to-heads: DeepSeek V4 Pro vs GPT-5.6 Sol · DeepSeek V4 Pro vs GPT-5.5 · DeepSeek V4 Pro vs GPT-5.3-Codex · DeepSeek V4 Pro vs GPT-5.6 Terra · DeepSeek V4 Pro vs Grok 4.5 · Claude Fable 5 vs DeepSeek V4 Pro

Full ranking: Best AI for research skills · model pages: DeepSeek V4 Pro, GPT-5.6 Luna