Best AI for writing

The verdict

ChatGPT

Scored 90/100 on our writing testsahead of Claude (87) and Gemini (87). You use it to answer questions, write text, create images, and write computer code. There is a free version, so trying it costs nothing.

Writing is the job most people actually use AI for — emails that need the right tone, documents that need structure, summaries that must not invent things. Generic benchmarks barely test it, which is why we run a dedicated writing suite: real workplace writing tasks, from a tricky client email to a plain-English rewrite of jargon. For a specific kind of writing, see the dedicated rankings: essays, emails, customer replies and study notes below.

updated 9 Aug 2026 · tested by Robert Prime · re-ranks automatically when a new run lands

How this was measured
  • All 3 sat the identical 18-task suite — same tasks, same order, one attempt each.
  • Every answer marked blind: judges are not told which entrant wrote it.
  • Three judges per answer, each from a competing lab. A panel never includes the entrant's own lab.
  • Scored 0–10 against a fixed rubric. Answers breaking a task's explicit rules are capped by machine, not by opinion.
  • Ties broken by fewest rule breaches, then lowest measured cost — published, not editorial.
  • Nobody pays for placement. Last computed 9 Aug 2026.

Every score below links to the raw file behind it: every task, the entrant’s real answer, and all three judges’ marks.

#ModelOur score
1ChatGPT90/100
2Claude87/100
3Gemini87/100

“API cost” is what software developers pay to build on a model — ignore it if you just use the website. Each model answers each task once. Models level on score are ranked by a fixed tie-break — fewest machine-checked rule breaches, then lowest measured cost per run — so the order is deterministic and checkable, never arbitrary. Judge panels never include the contestant’s own lab, so panels differ slightly per model — small cross-model gaps can reflect panel severity, not quality.

1.

ChatGPT

90/100our pick

You use it to answer questions, write text, create images, and write computer code. Best for anyone who needs to write, create images, or write computer code.

Strongest showing: Write to a deadline word count” — scored 10/10 by the panel. Weakest: “Product description” at 5/10. On 2 tasks a judge flagged a possible instruction breach — advisory only: the machine check did not confirm it and the panel outvoted it.

Meets all constraints perfectly, including the 90-110 word count (97 words). The content is highly accurate, actionable, and perfectly tailored for a professional LinkedIn audience.google/gemini-3.1-pro-preview, judging blind · full receipts ↓

18 tasks · 9 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · paid plans from £7/mo · full model page →

2.

Claude

87/100

It helps you analyze data, write code, and think through your hardest work. Best for problem solvers.

Strongest showing: Product description” — scored 9/10 by the panel. Weakest: “Voice match” at 5/10. Our automated rule-check caught 1 instruction breach across the run (a word limit or banned phrase) — capped in code, separately from the judges’ marks.

Meets 60-80 words, covers all specs for gym audience, avoids bans. Minor title addition; highly clear and useful.x-ai/grok-4.5, judging blind · full receipts ↓

18 tasks · 9 Aug 2026 · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5 · paid plans from $20/mo · full model page →

3.

Gemini

87/100

It helps you with writing, planning, and brainstorming. Best for anyone who needs help writing, planning, or brainstorming.

Strongest showing: Product description” — scored 9/10 by the panel. Weakest: “Two audiences” at 7/10. Our automated rule-check caught 1 instruction breach across the run (a word limit or banned phrase) — capped in code, separately from the judges’ marks.

Meets word count (74), avoids banned phrases, covers all specs, targets gym audience clearly, concise and engaging.anthropic/claude-sonnet-5, judging blind · full receipts ↓

18 tasks · 9 Aug 2026 · judges claude-sonnet-5, gpt-5.6-terra, grok-4.5 · paid plans from £7.99/mo · full model page →

See it for yourself

One task from the suite, and what the leader, the middle of the table and the bottom actually wrote. This is the task they disagreed on most — scores ranged from 5 to 10 out of 10.

The task — Voice match

Here is a brand's voice: short sentences, dry humour, never exclamation marks, addresses the reader as 'you', British spelling. Write a 70-word homepage intro in that voice for a company that repairs vintage watches.

ChatGPT10/10
Your vintage watch has seen things. Tea stains. Train stations. Possibly the moon, if your uncle was optimistic. We repair mechanical watches with patience, proper tools and a healthy suspicion of quick fixes. From tired movements to cracked crystals, we bring old timepieces back to work. You get ho…

Judge: Hits exactly 70 words. Perfectly captures the dry, short-sentence brand voice without exclamation marks. Uses 'you' and maintains British English conventions. Excellent, engaging copy.

Claude5/10
Your watch has seen things. Decades of wrist, weather and the occasional bad decision. It deserves better than a drawer.

We repair vintage watches with the care they were built for. No shortcuts, no guesswork, no rushed jobs to hit a deadline. Just proper craftsmanship, applied slowly.

Bring us th…

Judge: Flawlessly captures the requested dry, short-sentenced brand voice without exclamation marks. However, it fails the exact word count constraint, coming in at 67 words instead of 70.

Gemini9/10
Your old watch is broken. Again. You probably overwound it. We can fix that. We specialise in repairing vintage timepieces. We do not judge your clumsy fingers. We just replace the tiny gears you ruined. It takes time. Fortunately, we have plenty of it. Send us your broken artefact. We will return i…

Judge: Matches voice, dry humour, no exclamation marks, British spelling, exactly 70 words. Strong, minor stylistic nitpicks only.

Trimmed to fit. Every full answer, every judge’s score: ChatGPT, Claude, Gemini

How this ranking is made

Every model answers the identical published task suite. Three AI judges from three different labs mark each answer blind against a fixed 0–10 rubric — none is ever from the maker’s own lab, and an answer that breaks a task’s explicit rules is capped in code, not by the judge’s goodwill. The score is the average, out of 100.

What this is not: an opinion piece, a paid ranking, or a benchmark we cannot show you. No vendor can buy inclusion, a position or a score on this page — the order is computed from the test results before any link to a product exists, and where a link earns us a commission it says so on the link itself and the order is identical either way. Every score links its raw outputs and judge verdicts. The full protocol · How we make money · receipts: ChatGPT, Claude, Gemini

Questions people ask

What is the best AI for writing in 2026?

ChatGPT leads our tested ranking with 90/100 on our writing tests, ahead of Claude on 87. Every answer was marked blind by three AI judges from three different labs, and the full outputs are downloadable.

How is this ranking made?

Each model answers the identical published task suite; three judges from different labs score every answer 0–10 against a fixed rubric without knowing which produced it; answers that break a task's explicit rules are capped automatically. The score is the average, out of 100. No vendor pays for placement.

What happens when two models score the same?

They are separated by a fixed, published tie-break rather than by editorial choice: first the fewest machine-checked rule breaches, then the lowest measured cost per run. The order is therefore reproducible from the same data, and where the tie-break decides the top place the page says so.

How often does this page update?

It re-ranks itself whenever a new test run lands, and prices re-verify daily against vendor pages. The current ranking was last computed on 9 Aug 2026.

Is ChatGPT free?

Yes — ChatGPT publishes a free version on its own pricing page, with paid plans from £7/month.

Head-to-head battles

Worth knowing: Gemini finished ahead of Claude head-to-head despite ranking below it above; Claude finished ahead of ChatGPT head-to-head despite ranking below it above. The ranking scores each one alone against a fixed rubric; a battle puts them on the same task and asks which answer is better. When a gap this small flips, the honest read is that they are hard to separate — not that one of the numbers is wrong.

More rankings ▾

Best AI chatbot for everyday use · Best AI for coding · Best free AI model · Best AI for spreadsheets and Excel · Best AI essay writer · Best AI for summarising documents · Best AI for extracting data from text · Best AI for writing emails · Best AI for everyday maths and percentages · Best AI for customer service replies · Best AI for revision and study notes · Best AI for vibe coding · Best AI for making flashcards · Best AI for social media posts · Best AI for job applications and cover letters · Best AI for presentations · Best AI for writing your CV · Best AI for research skills · Best AI for creative writing · Best AI for translation · Best AI for travel planning · Best AI for emotional support · Best AI for everyday legal questions · Best AI for health questions · Best AI for writing a book · Best AI for humanising AI text · Best AI research agent · Best AI for code review · Best AI for workflow automation · Best AI for bookkeeping · Best AI for HR and employment questions · Best AI for meeting notes · Best AI for property and lettings · Best-value AI model API · every model we track · every tool