Best for / Essays
Best AI for writing essays
Essays need three things generic benchmarks never test: a real argument, obedience to a word count, and the discipline not to pad. So the suite asks for a 350-word argued case, an exact 100-word answer that must state its own count, a steelman of a position, and a rewrite of deliberately flabby prose. One task asks the model to invent a citation — the right answer is to refuse.
updated 13 Aug 2026 · tested by Robert Prime · re-ranks automatically when a new run lands
Scored 95/100 on our 12-task essay suite — in a dead heat with GPT-5.5 (94), so either is a fine choice.
| # | Model | Our score | Context | Free? | API $/1M in·out |
|---|---|---|---|---|---|
| 1 | GPT-5.6 Sollatest | 95/100 | 1050k | — | $5 · $30 |
| 2 | GPT-5.5 | 94/100 | 1050k | — | $5 · $30 |
| 3 | GPT-5.6 Lunalatest | 91/100 | 1050k | — | $0.1 · $0.6 |
| 4 | GPT-5.3-Codex | 88/100 | 400k | — | $1.75 · $14 |
| 5 | GPT-5.6 Terralatest | 88/100 | 1050k | — | $1 · $6 |
| 6 | Claude Sonnet 5latest | 85/100 | 1000k | — | $2 · $10 |
| 7 | Claude Opus 4.6 | 84/100 | 1000k | — | $5 · $25 |
| 8 | Claude Fable 5 | 83/100 | 1000k | — | $10 · $50 |
| 9 | Claude Opus 4.8 | 83/100 | 1000k | — | $5 · $25 |
| 10 | Gemini 3.1 Pro Preview | 83/100 | 1049k | — | $2 · $12 |
| 11 | Grok 4.5latest | 83/100 | 500k | grok.com | $2 · $6 |
| 12 | Qwen3.7 Maxlatest | 83/100 | 1000k | — | $1.475 · $4.425 |
| 13 | DeepSeek V4 Prolatest | 82/100 | 1049k | — | $1.168 · $2.336 |
| 14 | DeepSeek V4 Flashlatest | 81/100 | 1049k | — | $0.14 · $0.28 |
| 15 | GLM 5.2latest | 79/100 | 1049k | chat.z.ai | $0.5 · $1.98 |
| 16 | Gemini 3.1 Flash Lite | 78/100 | 1049k | — | $0.25 · $1.5 |
| 17 | Kimi K3latest | 78/100 | 1049k | — | $3 · $15 |
| 18 | Mistral Medium 3.5latest | 72/100 | 262k | — | $1.5 · $7.5 |
| 19 | Gemini 3.5 Flash | 71/100 | 1049k | Google AI Studio | $1.5 · $9 |
openai's newest model, holding about 1050k tokens of context — reached through the API.
Strongest showing: “Adapt for the reader” — scored 10/10 by the panel. Weakest: “Structured argument” at 8/10. The judges flagged 4 instruction breaches across the run — those scores were capped automatically.
“Both explanations accurate, within word limits, clearly labeled, and second adds new concepts (CPI, aggregate demand, expectations) not in first. Minor stylistic nitpicks only.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $5 in / $30 out per 1M tokens · full model page →
openai's model, holding about 1050k tokens of context — reached through the API.
Strongest showing: “Essay plan” — scored 10/10 by the panel. Weakest: “Structured argument” at 5/10. The judges flagged 4 instruction breaches across the run — those scores were capped automatically.
“Good structure and content, but exceeds 220-word limit (~200 words is close, actual count ~215-230, borderline); also thesis/sections slightly wordy, likely over limit given formatting.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $5 in / $30 out per 1M tokens · full model page →
openai's newest model, holding about 1050k tokens of context — reached through the API.
Strongest showing: “Essay plan” — scored 10/10 by the panel. Weakest: “Structured argument” at 5/10. The judges flagged 5 instruction breaches across the run — those scores were capped automatically.
“Good structure and content but exceeds 220-word limit (~250 words), violating explicit constraint; otherwise clear, useful, well-organized plan.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $0.1 in / $0.6 out per 1M tokens · full model page →
openai's model, holding about 400k tokens of context — reached through the API.
Strongest showing: “Tighten flabby prose” — scored 10/10 by the panel. Weakest: “Cite honestly” at 5/10. The judges flagged 6 instruction breaches across the run — those scores were capped automatically.
“Clear, concise, preserves meaning including hedging and timeframe. Minor loss of 'may possibly' redundancy but that's fine as it's simplification, not meaning loss.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $1.75 in / $14 out per 1M tokens · full model page →
openai's newest model, holding about 1050k tokens of context — reached through the API.
Strongest showing: “Tighten flabby prose” — scored 10/10 by the panel. Weakest: “Fix the structure” at 5/10. The judges flagged 9 instruction breaches across the run — those scores were capped automatically.
“Concise, accurate, preserves meaning including uncertainty and duration. Minor loss of hedging nuance but well within acceptable rewrite quality.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $1 in / $6 out per 1M tokens · full model page →
anthropic's newest model, holding about 1000k tokens of context — reached through the API.
Strongest showing: “Hold a word count” — scored 10/10 by the panel. Weakest: “Cite honestly” at 5/10. The judges flagged 8 instruction breaches across the run — those scores were capped automatically.
“The response contains 101 words in total (100 words for the description plus 1 word for the final number), violating the exact 100-word constraint.”— google/gemini-3.1-pro-preview, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5 · API $2 in / $10 out per 1M tokens · full model page →
anthropic's model, holding about 1000k tokens of context — reached through the API.
Strongest showing: “Essay plan” — scored 10/10 by the panel. Weakest: “Cite honestly” at 5/10. The judges flagged 9 instruction breaches across the run — those scores were capped automatically.
“Follows all instructions perfectly, including the word count limit, exact number of headings, and sentence constraints. The plan is highly relevant, clear, and well-structured for an undergraduate essay.”— google/gemini-3.1-pro-preview, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5 · API $5 in / $25 out per 1M tokens · full model page →
anthropic's model, holding about 1000k tokens of context — reached through the API.
Strongest showing: “Tighten flabby prose” — scored 10/10 by the panel. Weakest: “Cite honestly” at 5/10. The judges flagged 12 instruction breaches across the run — those scores were capped automatically.
“The response perfectly follows the instructions, providing a concise, clear rewrite without losing any meaning, and includes only the requested text.”— google/gemini-3.1-pro-preview, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5 · API $10 in / $50 out per 1M tokens · full model page →
anthropic's model, holding about 1000k tokens of context — reached through the API.
Strongest showing: “Tighten flabby prose” — scored 10/10 by the panel. Weakest: “Conclusion that adds something” at 5/10. The judges flagged 12 instruction breaches across the run — those scores were capped automatically.
“The response perfectly follows the instructions, providing a highly concise and clear rewrite that retains all original meaning while strictly adhering to the negative constraint of returning only the rewrite.”— google/gemini-3.1-pro-preview, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5 · API $5 in / $25 out per 1M tokens · full model page →
google's model, holding about 1049k tokens of context — reached through the API.
Strongest showing: “Essay plan” — scored 10/10 by the panel. Weakest: “Cite honestly” at 4/10. The judges flagged 7 instruction breaches across the run — those scores were capped automatically.
“Good structure and thesis, but exceeds 220-word limit (~230+ words) and thesis is two clauses, not strictly one clear sentence stance; otherwise meets format requirements well.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gpt-5.6-terra, grok-4.5 · API $2 in / $12 out per 1M tokens · full model page →
x-ai's newest model, holding about 500k tokens of context — reached through the API or free via grok.com.
Strongest showing: “Adapt for the reader” — scored 10/10 by the panel. Weakest: “Cite honestly” at 5/10. The judges flagged 5 instruction breaches across the run — those scores were capped automatically.
“Accurate, well-labeled, within word limits, second explanation adds new concepts (CPI, demand-pull, quantity theory, monetary policy) absent from first. Clear and concise.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $2 in / $6 out per 1M tokens · free via grok.com ↗ (checked 11 Aug 2026) · full model page →
qwen's newest model, holding about 1000k tokens of context — reached through the API.
Strongest showing: “Introduction only” — scored 10/10 by the panel. Weakest: “Structured argument” at 5/10. The judges flagged 3 instruction breaches across the run — those scores were capped automatically.
“Meets word limit, avoids forbidden openings, ends with clear thesis, well-written and relevant.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $1.475 in / $4.425 out per 1M tokens · full model page →
deepseek's newest model, holding about 1049k tokens of context — reached through the API.
Strongest showing: “Tighten flabby prose” — scored 10/10 by the panel. Weakest: “Cite honestly” at 5/10. The judges flagged 7 instruction breaches across the run — those scores were capped automatically.
“Concise, clear, preserves meaning including tentative correlation and duration. Minor loss of 'important to note' emphasis but negligible. Meets task well.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $1.168 in / $2.336 out per 1M tokens · full model page →
deepseek's newest model, holding about 1049k tokens of context — reached through the API.
Strongest showing: “Adapt for the reader” — scored 10/10 by the panel. Weakest: “Cite honestly” at 5/10. The judges flagged 6 instruction breaches across the run — those scores were capped automatically.
“Accurate, clear, meets word limits, labeled properly, and second explanation adds new concepts (CPI, demand-pull/cost-push, monetary policy) absent from the first.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $0.14 in / $0.28 out per 1M tokens · full model page →
z-ai's newest model, holding about 1049k tokens of context — reached through the API or free via chat.z.ai.
Strongest showing: “Adapt for the reader” — scored 10/10 by the panel. Weakest: “Cite honestly” at 3/10. The judges flagged 6 instruction breaches across the run — those scores were capped automatically.
“Accurate, clear, meets word limits, labeled correctly, second includes new concepts (CPI, demand-pull, monetary policy) absent from first.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $0.5 in / $1.98 out per 1M tokens · free via chat.z.ai ↗ (checked 11 Aug 2026) · full model page →
google's model, holding about 1049k tokens of context — reached through the API.
Strongest showing: “Adapt for the reader” — scored 10/10 by the panel. Weakest: “Cite honestly” at 5/10. The judges flagged 9 instruction breaches across the run — those scores were capped automatically.
“Both explanations meet word limits, are labeled correctly, accurate, and the second adds new concepts (inflationary expectations, CPI/PCE) absent from the first. Clear and well-structured.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gpt-5.6-terra, grok-4.5 · API $0.25 in / $1.5 out per 1M tokens · full model page →
moonshotai's newest model, holding about 1049k tokens of context — reached through the API.
Strongest showing: “Essay plan” — scored 10/10 by the panel. Weakest: “Conclusion that adds something” at 5/10. The judges flagged 9 instruction breaches across the run — those scores were capped automatically.
“Meets format requirements with thesis, five headings, evidence sentences; concise and relevant, though total word count exceeds 220 (approx. 250), a minor breach not flagged strictly since instructions say check literally.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $3 in / $15 out per 1M tokens · full model page →
mistralai's newest model, holding about 262k tokens of context — reached through the API.
Strongest showing: “Tighten flabby prose” — scored 10/10 by the panel. Weakest: “Cite honestly” at 4/10. The judges flagged 13 instruction breaches across the run — those scores were capped automatically.
“Concise, clear, preserves meaning including hedging and duration. Minor loss of 'note' framing but negligible.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $1.5 in / $7.5 out per 1M tokens · full model page →
google's model, holding about 1049k tokens of context — reached through the API or free via Google AI Studio.
Strongest showing: “Adapt for the reader” — scored 10/10 by the panel. Weakest: “Cite honestly” at 0/10. The judges flagged 10 instruction breaches across the run — those scores were capped automatically.
“Accurate, well-labeled, within word limits, includes advanced concepts absent from simple explanation. Clear and concise, useful for both audiences.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gpt-5.6-terra, grok-4.5 · API $1.5 in / $9 out per 1M tokens · free via Google AI Studio ↗ (checked 11 Aug 2026) · full model page →
See it for yourself
One task from the suite, and what the leader, the middle of the table and the bottom actually wrote. This is the task they disagreed on most — scores ranged from 0 to 10 out of 10.
Write a description of the water cycle in exactly 100 words. Count them. The final line of your reply must be the word count as a number and nothing else.
Water continually circulates between Earth’s surface and atmosphere through the water cycle. Sunlight heats oceans, lakes, rivers, and soil, causing liquid water to evaporate into vapor. Plants also release moisture through transpiration. As warm vapor rises, it cools and condenses around tiny parti…
Judge: Accurate, well-structured 100-word description confirmed by count; final line correctly shows 100, meeting all constraints clearly.
The water cycle is a continuous journey that sustains all life on Earth. It begins when the sun heats oceans, lakes, and rivers, causing liquid water to evaporate into an invisible gas called water vapor. As this warm vapor rises into the atmosphere, it cools and condenses to form fluffy clouds. Whe…
Judge: Word count verified exactly 100, final line shows number, content accurate and clear.
39: and
40: condenses
41: into
42: clouds.
43: When
44: these
45: droplets
46: grow
47: heavy,
48: they
49: fall
50: back
51: to
52: the
53: ground
54: as
55: precipitation,
56: including
57: rain,
58: snow,
59: or
6…Judge: Response is broken, incomplete, malformed word-count listing, not an actual description; fails task entirely.
Trimmed to fit. Every full answer, every judge’s score: GPT-5.6 Sol, Gemini 3.1 Pro Preview, Gemini 3.5 Flash
How this ranking is made
Every model answers the identical published task suite. Three AI judges from three different labs mark each answer blind against a fixed 0–10 rubric — none is ever from the maker’s own lab, and an answer that breaks a task’s explicit rules is capped in code, not by the judge’s goodwill. The score is the average, out of 100.
What this is not: an opinion piece, an affiliate list, or a benchmark we cannot show you. Every score links its raw outputs and judge verdicts. The full protocol · receipts: GPT-5.6 Sol, GPT-5.5, GPT-5.6 Luna, GPT-5.3-Codex, GPT-5.6 Terra, Claude Sonnet 5, Claude Opus 4.6, Claude Fable 5, Claude Opus 4.8, Gemini 3.1 Pro Preview, Grok 4.5, Qwen3.7 Max, DeepSeek V4 Pro, DeepSeek V4 Flash, GLM 5.2, Gemini 3.1 Flash Lite, Kimi K3, Mistral Medium 3.5, Gemini 3.5 Flash
Questions people ask
What is the best AI for writing essays in 2026?
GPT-5.6 Sol leads our tested ranking with 95/100 on our 12-task essay suite (12 tasks), in a dead heat with GPT-5.5 (94). Every answer was marked blind by three AI judges from three different labs, and the full outputs are downloadable.
How is this ranking made?
Each model answers the identical published task suite; three judges from different labs score every answer 0–10 against a fixed rubric without knowing which produced it; answers that break a task's explicit rules are capped automatically. The score is the average, out of 100. No vendor pays for placement.
How often does this page update?
It re-ranks itself whenever a new test run lands, and prices re-verify daily against vendor pages. The current ranking was last computed on 13 Aug 2026.
More rankings: Best AI for writing · Best AI chatbot for everyday use · Best AI for coding · Best free AI model · Best AI for spreadsheets and Excel · Best AI for summarising documents · Best AI for extracting data from text · Best-value AI model API · every model we track · every tool