Best for / Summarising
Best AI for summarising documents
The failure mode that matters in summarising is not dullness, it is quiet invention — a caveat dropped, a number changed, a decision asserted that nobody made. Every task here is scored on what survives: the limitation, the four figures, the disagreement stated fairly. One task is a two-line email with nothing to summarise, to see which models pad rather than say so.
updated 13 Aug 2026 · tested by Robert Prime · re-ranks automatically when a new run lands
Scored 99/100 on our 12-task summarising suite — in a dead heat with GPT-5.5 (98), so either is a fine choice.
| # | Model | Our score | Context | Free? | API $/1M in·out |
|---|---|---|---|---|---|
| 1 | Claude Opus 4.8 | 99/100 | 1000k | — | $5 · $25 |
| 2 | GPT-5.5 | 98/100 | 1050k | — | $5 · $30 |
| 3 | GPT-5.6 Lunalatest | 97/100 | 1050k | — | $0.1 · $0.6 |
| 4 | Claude Fable 5 | 96/100 | 1000k | — | $10 · $50 |
| 5 | GPT-5.3-Codex | 96/100 | 400k | — | $1.75 · $14 |
| 6 | GPT-5.6 Terralatest | 96/100 | 1050k | — | $1 · $6 |
| 7 | DeepSeek V4 Prolatest | 95/100 | 1049k | — | $1.168 · $2.336 |
| 8 | GPT-5.6 Sollatest | 95/100 | 1050k | — | $5 · $30 |
| 9 | Claude Opus 4.6 | 93/100 | 1000k | — | $5 · $25 |
| 10 | Claude Sonnet 5latest | 93/100 | 1000k | — | $2 · $10 |
| 11 | Grok 4.5latest | 92/100 | 500k | grok.com | $2 · $6 |
| 12 | GLM 5.2latest | 91/100 | 1049k | chat.z.ai | $0.5 · $1.98 |
| 13 | Gemini 3.1 Pro Preview | 90/100 | 1049k | — | $2 · $12 |
| 14 | Kimi K3latest | 90/100 | 1049k | — | $3 · $15 |
| 15 | Gemini 3.5 Flash | 89/100 | 1049k | Google AI Studio | $1.5 · $9 |
| 16 | Qwen3.7 Maxlatest | 87/100 | 1000k | — | $1.475 · $4.425 |
| 17 | Gemini 3.1 Flash Lite | 85/100 | 1049k | — | $0.25 · $1.5 |
| 18 | DeepSeek V4 Flashlatest | 83/100 | 1049k | — | $0.14 · $0.28 |
| 19 | Mistral Medium 3.5latest | 83/100 | 262k | — | $1.5 · $7.5 |
anthropic's model, holding about 1000k tokens of context — reached through the API.
Strongest showing: “Meeting notes to actions” — scored 10/10 by the panel. Weakest: “Summary for a specific reader” at 9/10. The judges flagged 1 instruction breach across the run — those scores were capped automatically.
“The response perfectly extracts all action items, assigns the correct owners (or 'unassigned'), formats them as a bulleted list, and includes no extraneous text, following all instructions flawlessly.”— google/gemini-3.1-pro-preview, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5 · API $5 in / $25 out per 1M tokens · full model page →
openai's model, holding about 1050k tokens of context — reached through the API.
Strongest showing: “Five-bullet summary” — scored 10/10 by the panel. Weakest: “One-sentence gist” at 9/10.
“Accurate, concise, meets 5-bullet/15-word constraints, though last bullet merges two facts slightly awkwardly.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $5 in / $30 out per 1M tokens · full model page →
openai's newest model, holding about 1050k tokens of context — reached through the API.
Strongest showing: “Meeting notes to actions” — scored 10/10 by the panel. Weakest: “Summary for a specific reader” at 9/10. The judges flagged 1 instruction breach across the run — those scores were capped automatically.
“Accurate, well-formatted action items with owners correctly assigned or marked unassigned; Sam's leave correctly omitted as non-action. Minor: could omit conditional caveat but still concise and useful.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $0.1 in / $0.6 out per 1M tokens · full model page →
anthropic's model, holding about 1000k tokens of context — reached through the API.
Strongest showing: “Meeting notes to actions” — scored 10/10 by the panel. Weakest: “Long to short, no loss” at 8/10. The judges flagged 1 instruction breach across the run — those scores were capped automatically.
“The response perfectly extracts all action items, assigns the correct owners (including 'unassigned'), and formats them as a bulleted list containing only the requested information.”— google/gemini-3.1-pro-preview, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5 · API $10 in / $50 out per 1M tokens · full model page →
openai's model, holding about 400k tokens of context — reached through the API.
Strongest showing: “Do not invent” — scored 10/10 by the panel. Weakest: “Five-bullet summary” at 8/10.
“Accurate, concise summary under 50 words, correctly follows format with 'Not stated' line, no invented facts. Very minor: could be slightly more concise but meets requirements well.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $1.75 in / $14 out per 1M tokens · full model page →
openai's newest model, holding about 1050k tokens of context — reached through the API.
Strongest showing: “Five-bullet summary” — scored 10/10 by the panel. Weakest: “One-sentence gist” at 8/10.
“Accurate, concise, exactly 5 bullets, all under 15 words, covers key facts clearly and readable quickly.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $1 in / $6 out per 1M tokens · full model page →
deepseek's newest model, holding about 1049k tokens of context — reached through the API.
Strongest showing: “Meeting notes to actions” — scored 10/10 by the panel. Weakest: “Headline and standfirst” at 8/10.
“Accurate, follows format, captures all action items with owners; minor nuance on Priya's conditional commitment omitted but acceptable.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $1.168 in / $2.336 out per 1M tokens · full model page →
openai's newest model, holding about 1050k tokens of context — reached through the API.
Strongest showing: “Do not invent” — scored 10/10 by the panel. Weakest: “Meeting notes to actions” at 6/10.
“Accurate, concise summary under 50 words, follows format exactly, plausible 'not stated' fact. Minor: could be slightly more informative but meets task well.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $5 in / $30 out per 1M tokens · full model page →
anthropic's model, holding about 1000k tokens of context — reached through the API.
Strongest showing: “Meeting notes to actions” — scored 10/10 by the panel. Weakest: “Summarise a disagreement” at 6/10. The judges flagged 1 instruction breach across the run — those scores were capped automatically.
“The response perfectly extracts the action items and their owners, correctly identifies the unassigned task, and formats them as a bulleted list exactly as requested.”— google/gemini-3.1-pro-preview, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5 · API $5 in / $25 out per 1M tokens · full model page →
anthropic's newest model, holding about 1000k tokens of context — reached through the API.
Strongest showing: “Five-bullet summary” — scored 10/10 by the panel. Weakest: “Extract the decision” at 5/10. The judges flagged 4 instruction breaches across the run — those scores were capped automatically.
“The response perfectly follows all instructions, including the exact bullet count and word limit per bullet. It accurately and concisely captures all key information, making it highly useful for a busy professional.”— google/gemini-3.1-pro-preview, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5 · API $2 in / $10 out per 1M tokens · full model page →
x-ai's newest model, holding about 500k tokens of context — reached through the API or free via grok.com.
Strongest showing: “Meeting notes to actions” — scored 10/10 by the panel. Weakest: “Five-bullet summary” at 6/10. The judges flagged 1 instruction breach across the run — those scores were capped automatically.
“Accurate, follows format, complete action items with owners correctly assigned; minor debatable inclusion but overall solid and concise.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $2 in / $6 out per 1M tokens · free via grok.com ↗ (checked 11 Aug 2026) · full model page →
z-ai's newest model, holding about 1049k tokens of context — reached through the API or free via chat.z.ai.
Strongest showing: “Do not invent” — scored 10/10 by the panel. Weakest: “Meeting notes to actions” at 5/10.
“Accurate, concise summary under 50 words, correct format, reasonable 'not stated' item. Minor: could be more concise but meets constraints well.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $0.5 in / $1.98 out per 1M tokens · free via chat.z.ai ↗ (checked 11 Aug 2026) · full model page →
google's model, holding about 1049k tokens of context — reached through the API.
Strongest showing: “Do not invent” — scored 10/10 by the panel. Weakest: “Headline and standfirst” at 8/10.
“Accurate, concise, under 50 words, follows format with 'Not stated' line and valid unmentioned fact.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gpt-5.6-terra, grok-4.5 · API $2 in / $12 out per 1M tokens · full model page →
moonshotai's newest model, holding about 1049k tokens of context — reached through the API.
Strongest showing: “Five-bullet summary” — scored 10/10 by the panel. Weakest: “Headline and standfirst” at 8/10. The judges flagged 1 instruction breach across the run — those scores were capped automatically.
“Accurate, exactly 5 bullets, all under 15 words, clear and concise, covers all key facts.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $3 in / $15 out per 1M tokens · full model page →
google's model, holding about 1049k tokens of context — reached through the API or free via Google AI Studio.
Strongest showing: “Do not invent” — scored 10/10 by the panel. Weakest: “Meeting notes to actions” at 6/10. The judges flagged 2 instruction breaches across the run — those scores were capped automatically.
“Accurate, concise summary under 50 words, includes required 'Not stated' line with reasonable unstated fact. Minor stylistic nitpick only.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gpt-5.6-terra, grok-4.5 · API $1.5 in / $9 out per 1M tokens · free via Google AI Studio ↗ (checked 11 Aug 2026) · full model page →
qwen's newest model, holding about 1000k tokens of context — reached through the API.
Strongest showing: “Meeting notes to actions” — scored 10/10 by the panel. Weakest: “Five-bullet summary” at 3/10. The judges flagged 1 instruction breach across the run — those scores were capped automatically.
“Accurate, follows format, concise; minor nuance missed (Priya's conditional 'if she has the numbers') but overall correct and useful.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $1.475 in / $4.425 out per 1M tokens · full model page →
google's model, holding about 1049k tokens of context — reached through the API.
Strongest showing: “Summarise a disagreement” — scored 10/10 by the panel. Weakest: “Extract the decision” at 6/10.
“Accurate, balanced, but exceeds 80 words (approx 70... actually count: ~70 words, within limit). Clear and fair, minor verbosity.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gpt-5.6-terra, grok-4.5 · API $0.25 in / $1.5 out per 1M tokens · full model page →
deepseek's newest model, holding about 1049k tokens of context — reached through the API.
Strongest showing: “Five-bullet summary” — scored 10/10 by the panel. Weakest: “Headline and standfirst” at 6/10.
“Accurate, concise, follows 5-bullet format, though last bullet merges two facts, slightly overloaded.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $0.14 in / $0.28 out per 1M tokens · full model page →
mistralai's newest model, holding about 262k tokens of context — reached through the API.
Strongest showing: “Meeting notes to actions” — scored 10/10 by the panel. Weakest: “Extract the decision” at 5/10. The judges flagged 2 instruction breaches across the run — those scores were capped automatically.
“Accurate, well-formatted list matching notes; correctly assigns owners/unassigned. Minor debatable point: Priya's cashflow task was conditional, but reasonably represented.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $1.5 in / $7.5 out per 1M tokens · full model page →
See it for yourself
One task from the suite, and what the leader, the middle of the table and the bottom actually wrote. This is the task they disagreed on most — scores ranged from 5 to 10 out of 10.
From this thread, state in one line WHAT was decided and WHO decided it. If no decision was reached, say so. Thread: 'A: shall we move the launch to October? B: I'd rather not, the trade show is in September. A: agreed, but we can't ship without the packaging. C: packaging lands 2 Sept, confirmed by the printer this morning. A: fine, September stands then.'
Decision: The launch stays in September (not moved to October), decided by A (with input from B and C).
Judge: The response perfectly follows the instructions, stating clearly in one line what was decided and who decided it based on the provided thread.
**Decision:** The launch date will remain in September (not moved to October). **Decided by:** A (with agreement from B and confirmation from C on packaging timing).
Judge: Accurately captures the decision and the decision-maker, but fails the explicit constraint to state it in exactly one line.
**WHAT was decided:** The launch remains in September. **WHO decided it:** A, B, and C (collectively).
Judge: Reasonable summary but not one line as instructed; format uses two lines/bolded headers instead of a single line.
Trimmed to fit. Every full answer, every judge’s score: Claude Opus 4.8, Claude Sonnet 5, Mistral Medium 3.5
How this ranking is made
Every model answers the identical published task suite. Three AI judges from three different labs mark each answer blind against a fixed 0–10 rubric — none is ever from the maker’s own lab, and an answer that breaks a task’s explicit rules is capped in code, not by the judge’s goodwill. The score is the average, out of 100.
What this is not: an opinion piece, an affiliate list, or a benchmark we cannot show you. Every score links its raw outputs and judge verdicts. The full protocol · receipts: Claude Opus 4.8, GPT-5.5, GPT-5.6 Luna, Claude Fable 5, GPT-5.3-Codex, GPT-5.6 Terra, DeepSeek V4 Pro, GPT-5.6 Sol, Claude Opus 4.6, Claude Sonnet 5, Grok 4.5, GLM 5.2, Gemini 3.1 Pro Preview, Kimi K3, Gemini 3.5 Flash, Qwen3.7 Max, Gemini 3.1 Flash Lite, DeepSeek V4 Flash, Mistral Medium 3.5
Questions people ask
What is the best AI for summarising documents in 2026?
Claude Opus 4.8 leads our tested ranking with 99/100 on our 12-task summarising suite (12 tasks), in a dead heat with GPT-5.5 (98). Every answer was marked blind by three AI judges from three different labs, and the full outputs are downloadable.
How is this ranking made?
Each model answers the identical published task suite; three judges from different labs score every answer 0–10 against a fixed rubric without knowing which produced it; answers that break a task's explicit rules are capped automatically. The score is the average, out of 100. No vendor pays for placement.
How often does this page update?
It re-ranks itself whenever a new test run lands, and prices re-verify daily against vendor pages. The current ranking was last computed on 13 Aug 2026.
More rankings: Best AI for writing · Best AI chatbot for everyday use · Best AI for coding · Best free AI model · Best AI for spreadsheets and Excel · Best AI for writing essays · Best AI for extracting data from text · Best-value AI model API · every model we track · every tool