DeepSeek V4 Pro vs Grok 4.5 for research skills
DeepSeek V4 Pro wins, 98–93
Same 12 research skills tasks, marked blind by three rival labs. DeepSeek V4 Pro took 4 tasks, Grok 4.5 took 0, 8 tied. Tested 16 Aug 2026.
Where they differed most
Three studies, summarised: Study A (2021, 12,000 adults followed 5 years) — people sleeping under 6 hours had 20% higher rates of catching colds. Study B (2019, 300 students, one term) — no significant link between sleep and illness days. Study C (2023, review of 40 studies) — consistent link between short sleep and infection risk, but most studies could not rule out that stress causes both. Write a summary of max 120 words of what we actually know, giving the disagreement and the confounder their honest weight.
Evidence on sleep and infection risk is mixed. A large five-year study found adults sleeping under six hours had 20% higher cold rates, but a smaller student study found no significant link between sleep and illness days. A 2023 review of 40 studies reports a consistent association between short sleep and infection risk, yet most underlying studies could not rule out that stress drives both poor sleep and greater susceptibility. Overall, short sleep appears linked to higher infection risk, but the evidence is observational and confounded; poor…
Research shows mixed evidence on short sleep and infection risk. A large 2021 study (12,000 adults, 5 years) linked <6 hours’ sleep to 20% higher cold rates. A 2023 review of 40 studies found a consistent association with infection risk. Yet a 2019 study (300 students) detected no significant link to illness days. Crucially, most research could not exclude stress as a confounder that may cause both poor sleep and greater susceptibility. An association appears in bigger/broader data, but disagreement exists and causality is unproven. (Word cou…
Task by task
| Task | DeepSeek V4 Pro | Grok 4.5 |
|---|---|---|
| Make it answerable | 10 | 10 |
| Rank the sources | 9 | 9 |
| Synthesise the disagreement | 10 | 8 |
| Tear down the survey | 9 | 9 |
| Correlation for the neighbour | 10 | 10 |
| Plan the dig | 9.5 | 8 |
| What you cannot cite | 10 | 9 |
| Structure the evidence | 10 | 10 |
| Ask without leading | 10 | 10 |
| Two sources, head on | 10 | 10 |
| Primary or secondary | 10 | 8 |
| Refuse the invented bibliography | 10 | 10 |
Full receipts: DeepSeek V4 Pro, Grok 4.5 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra
Questions people ask
Which is better for research skills: DeepSeek V4 Pro or Grok 4.5?
DeepSeek V4 Pro — it scored 98/100 against 93/100 on our 12-task research skills suite, winning 4 tasks to 0 with 8 tied. Every answer was marked blind by three judges from three rival AI labs.
How was this tested?
Both models answered the identical published research skills tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.
More research skills head-to-heads: DeepSeek V4 Pro vs GPT-5.6 Sol · DeepSeek V4 Pro vs GPT-5.5 · DeepSeek V4 Pro vs GPT-5.6 Luna · DeepSeek V4 Pro vs GPT-5.3-Codex · DeepSeek V4 Pro vs GPT-5.6 Terra · Claude Fable 5 vs DeepSeek V4 Pro
Full ranking: Best AI for research skills · model pages: DeepSeek V4 Pro, Grok 4.5