Best AI for research skills
Scored 98/100 on our 12-task research suite — level on quality with GPT-5.6 Sol (98) — the top spot goes on the tie-break: cleanest rule-compliance, then lowest measured cost; GPT-5.5 sits a single point behind.
DeepSeek V4 Pro is the engine inside DeepSeek. Go to chat.deepseek.com ↗. Not fussed about the last point or two? Any of the top 3 here will serve you well.
No live browsing — every task carries its material with it, because the skill tested is judgement, not retrieval: ranking five sources for trustworthiness, synthesising three studies that disagree, tearing down the 40-person gym survey behind a '92% of Britons' headline, and writing survey questions both sides would call fair. One task demands a DOI the model cannot verify — the honest answer says so plainly; another asks for 15 invented citations, and the right answer is to refuse.
updated 16 Aug 2026 · tested by Robert Prime · re-ranks automatically when a new run lands
| # | Model | Our score |
|---|---|---|
| 1 | DeepSeek V4 Prolatest | 98/100 |
| 2 | GPT-5.6 Sollatest | 98/100 |
| 3 | GPT-5.5 | 97/100 |
| 4 | GPT-5.6 Luna | 96/100 |
| 5 | GPT-5.3-Codex | 95/100 |
| 6 | GPT-5.6 Terra | 94/100 |
| 7 | Grok 4.5 | 93/100 |
| 8 | Claude Fable 5 | 92/100 |
| 9 | Claude Opus 4.8 | 90/100 |
| 10 | Kimi K3 | 90/100 |
| 11 | DeepSeek V4 Flash | 89/100 |
| 12 | Gemini 3.1 Flash Lite | 88/100 |
| 13 | Gemini 3.5 Flash | 88/100 |
| 14 | Qwen3.7 Max | 88/100 |
| 15 | GLM 5.2 | 86/100 |
| 16 | Gemini 3.1 Pro Preview | 85/100 |
| 17 | Claude Sonnet 5 | 84/100 |
| 18 | Mistral Medium 3.5 | 83/100 |
| 19 | Claude Opus 4.6 | 83/100 |
“API cost” is what software developers pay to build on a model — ignore it if you just use the website. Each model answers each task once. Models level on score are ranked by a fixed tie-break — fewest machine-checked rule breaches, then lowest measured cost per run — so the order is deterministic and checkable, never arbitrary. Judge panels never include the contestant’s own lab, so panels differ slightly per model — small cross-model gaps can reflect panel severity, not quality.
Made by DeepSeek — their newest model. You use it inside DeepSeek — nothing to install.
Strongest showing: “Synthesise the disagreement” — scored 10/10 by the panel.
“The response perfectly synthesizes the three studies within the 120-word limit (98 words). It accurately highlights the disagreement between the studies and appropriately emphasizes the confounding role of stress, making it highly useful an”— google/gemini-3.1-pro-preview, judging blind · full receipts ↓
12 tasks · 16 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $1.168 in / $2.336 out per 1M tokens · full model page →
Made by OpenAI — their newest model. You use it inside ChatGPT — nothing to install.
Strongest showing: “Plan the dig” — scored 10/10 by the panel. Weakest: “Two sources, head on” at 9/10.
“Flawless response. It perfectly addresses all prompts, accurately details UK/Leicester-specific sources, explains their limitations, provides a logical workflow, and strictly adheres to the no-URL constraint.”— google/gemini-3.1-pro-preview, judging blind · full receipts ↓
12 tasks · 16 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $5 in / $30 out per 1M tokens · full model page →
Made by OpenAI. You use it inside ChatGPT — nothing to install.
Strongest showing: “Primary or secondary” — scored 10/10 by the panel. Weakest: “Correlation for the neighbour” at 9/10.
“The response perfectly follows all instructions, accurately classifying each source with a concise one-line justification. It correctly identifies a primary source needing caution and provides a valid historical reason. Flawless execution.”— google/gemini-3.1-pro-preview, judging blind · full receipts ↓
12 tasks · 16 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $5 in / $30 out per 1M tokens · full model page →
See it for yourself
One task from the suite, and what the leader, the middle of the table and the bottom actually wrote. This is the task they disagreed on most — scores ranged from 5 to 10 out of 10.
Someone wants to trace the history of their 1930s semi in Leicester — who built it, who has lived there, what stood on the land before. Lay out the kinds of sources and places in the UK that answer each part (deeds, census-type records, old maps, local archives, newspaper archives), what each one can and cannot tell them, and the sensible order to work in. No URLs — the shape of the search, not links.
Trace the history of a 1930s semi-detached house in Leicester by working from the present backwards. The key local institutions are the **Record Office for Leicestershire, Leicester and Rutland**, **Leicester Local Studies Library**, **Leicester City Council planning/building control**, and **HM Lan…
Judge: The response perfectly addresses all aspects of the prompt, providing accurate, well-structured, and highly useful information for a non-technical reader. It strictly adheres to the negative constraint of including no URLs.
# Tracing a 1930s semi in Leicester Good news first: a 1930s house sits in a sweet spot. It's recent enough that paper survives in quantity (deeds, building control plans, the 1939 Register), but old enough that the land underneath was almost certainly fields, so there's a "before" worth finding. O…
Judge: Highly accurate and specific to Leicester, but fails to explicitly outline the requested 'sensible order to work in' and cuts off mid-sentence at the very end.
# Tracing the History of a 1930s Semi in Leicester ## The Shape of the Search Before diving into sources, it helps to understand that house history works backwards. You start from what you know — the present — and peel back layers. Each source tends to cover a different slice of time or answer a d…
Judge: The response perfectly addresses all prompt requirements. It accurately details UK-specific sources, explicitly covers what they can and cannot reveal, provides a logical search order, and strictly adheres to the negative constraint by omit
Trimmed to fit. Every full answer, every judge’s score: DeepSeek V4 Pro, Kimi K3, Claude Opus 4.6
How this ranking is made
Every model answers the identical published task suite. Three AI judges from three different labs mark each answer blind against a fixed 0–10 rubric — none is ever from the maker’s own lab, and an answer that breaks a task’s explicit rules is capped in code, not by the judge’s goodwill. The score is the average, out of 100.
What this is not: an opinion piece, an affiliate list, or a benchmark we cannot show you. Every score links its raw outputs and judge verdicts. The full protocol · receipts: DeepSeek V4 Pro, GPT-5.6 Sol, GPT-5.5, GPT-5.6 Luna, GPT-5.3-Codex, GPT-5.6 Terra, Grok 4.5, Claude Fable 5, Claude Opus 4.8, Kimi K3, DeepSeek V4 Flash, Gemini 3.1 Flash Lite, Gemini 3.5 Flash, Qwen3.7 Max, GLM 5.2, Gemini 3.1 Pro Preview, Claude Sonnet 5, Mistral Medium 3.5, Claude Opus 4.6
Questions people ask
What is the best AI for research skills in 2026?
DeepSeek V4 Pro leads our tested ranking with 98/100 on our 12-task research suite (12 tasks), in a dead heat with GPT-5.6 Sol (98). Every answer was marked blind by three AI judges from three different labs, and the full outputs are downloadable.
How is this ranking made?
Each model answers the identical published task suite; three judges from different labs score every answer 0–10 against a fixed rubric without knowing which produced it; answers that break a task's explicit rules are capped automatically. The score is the average, out of 100. No vendor pays for placement.
How often does this page update?
It re-ranks itself whenever a new test run lands, and prices re-verify daily against vendor pages. The current ranking was last computed on 16 Aug 2026.
Head-to-heads in research
Show all 20 tested pairs ▾
More rankings ▾
Best AI for writing · Best AI chatbot for everyday use · Best AI for coding · Best free AI model · Best AI for spreadsheets and Excel · Best AI essay writer · Best AI for summarising documents · Best AI for extracting data from text · Best AI for writing emails · Best AI for everyday maths and percentages · Best AI for customer service replies · Best AI for revision and study notes · Best AI for vibe coding · Best AI for making flashcards · Best AI for social media posts · Best AI for job applications and cover letters · Best AI for presentations · Best AI for writing your CV · Best AI for creative writing · Best AI for translation · Best AI for travel planning · Best AI for emotional support · Best AI for everyday legal questions · Best AI for health questions · Best AI for writing a book · Best AI for humanising AI text · Best-value AI model API · every model we track · every tool