DeepSeek V4 Pro vs GPT-5.3-Codex for research skills

The verdict

DeepSeek V4 Pro wins, 9895

Same 12 research skills tasks, marked blind by three rival labs. DeepSeek V4 Pro took 5 tasks, GPT-5.3-Codex took 1, 6 tied. Tested 16 Aug 2026.

Where they differed most

The task — Rank the sources

You are researching whether standing desks reduce back pain. Rank these 5 sources from most to least trustworthy for that question, one line of reasoning each: (a) a 2022 systematic review of 17 trials in a peer-reviewed journal; (b) a standing-desk manufacturer's white paper; (c) a Mumsnet thread with 200 replies; (d) a 2015 trial with 24 office workers; (e) an NHS page on back pain last reviewed 2024. Then say which ONE you would cite to a sceptical reader and why.

Ranking from most to least trustworthy:

1. **(a) 2022 systematic review of 17 trials in a peer-reviewed journal** — highest-quality evidence because it synthesises multiple controlled trials and has independent peer review.  
2. **(e) NHS page on back pain last reviewed 2024** — authoritative and current, but likely general back-pain guidance rather than a focused evaluation of standing desks.  
3. **(d) 2015 trial with 24 office workers** — direct experimental evidence, but small, dated, and limited by sample size.  
4. **(b) standing-desk ma…
**Most → least trustworthy for “do standing desks reduce back pain?”**

1. **(a) 2022 systematic review of 17 trials** — Highest level of evidence here: it pools multiple studies, is peer-reviewed, and is recent.  
2. **(e) NHS back pain page (reviewed 2024)** — Very credible institution and up to date, but likely broader guidance rather than direct trial-level evidence on standing desks specifically.  
3. **(d) 2015 trial with 24 office workers** — Primary research and directly relevant, but small sample and older single-study design make find…

Task by task

TaskDeepSeek V4 ProGPT-5.3-Codex
Make it answerable1010
Rank the sources910
Synthesise the disagreement109
Tear down the survey99
Correlation for the neighbour109
Plan the dig9.59
What you cannot cite1010
Structure the evidence1010
Ask without leading109
Two sources, head on1010
Primary or secondary109
Refuse the invented bibliography1010

Full receipts: DeepSeek V4 Pro, GPT-5.3-Codex · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra

Questions people ask

Which is better for research skills: DeepSeek V4 Pro or GPT-5.3-Codex?

DeepSeek V4 Pro — it scored 98/100 against 95/100 on our 12-task research skills suite, winning 5 tasks to 1 with 6 tied. Every answer was marked blind by three judges from three rival AI labs.

How was this tested?

Both models answered the identical published research skills tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.

More research skills head-to-heads: DeepSeek V4 Pro vs GPT-5.6 Sol · DeepSeek V4 Pro vs GPT-5.5 · DeepSeek V4 Pro vs GPT-5.6 Luna · DeepSeek V4 Pro vs GPT-5.6 Terra · DeepSeek V4 Pro vs Grok 4.5 · Claude Fable 5 vs DeepSeek V4 Pro

Full ranking: Best AI for research skills · model pages: DeepSeek V4 Pro, GPT-5.3-Codex