Best AI for creative writing
Scored 98/100 on our 12-task creative-writing suite — a single point ahead of GPT-5.6 Terra (97) — effectively level.
GPT-5.5 is the engine inside ChatGPT. Go to chatgpt.com ↗ — the free tier is fine to start. Paid plans start at £7/month (go plan, vendor’s own price). The free tier is ChatGPT’s, not a promise about this exact model — we haven’t verified which plan carries it. Not fussed about the last point or two? Any of the top 3 here will serve you well.
Anyone can generate words; this suite tests control: anger shown without the word 'angry' (the ban is enforced in code), a story of exactly 100 words, purple prose cut to the one image that earns its place, a narrator's voice continued without drift, and an unreliable narrator the reader sees through. One task asks for a realistic fake news article about a named neighbour — the right answer is to refuse.
updated 16 Aug 2026 · tested by Robert Prime · re-ranks automatically when a new run lands
| # | Model | Our score |
|---|---|---|
| 1 | GPT-5.5 | 98/100 |
| 2 | GPT-5.6 Terralatest | 97/100 |
| 3 | GPT-5.6 Sollatest | 97/100 |
| 4 | Gemini 3.5 Flash | 96/100 |
| 5 | Qwen3.7 Max | 96/100 |
| 6 | Gemini 3.1 Pro Preview | 95/100 |
| 7 | GPT-5.6 Luna | 94/100 |
| 8 | GPT-5.3-Codex | 94/100 |
| 9 | DeepSeek V4 Pro | 93/100 |
| 10 | Grok 4.5 | 93/100 |
| 11 | Claude Sonnet 5 | 93/100 |
| 12 | Claude Opus 4.8 | 91/100 |
| 13 | Claude Fable 5 | 91/100 |
| 14 | Kimi K3 | 89/100 |
| 15 | Claude Opus 4.6 | 87/100 |
| 16 | GLM 5.2 | 86/100 |
| 17 | DeepSeek V4 Flash | 83/100 |
| 18 | Mistral Medium 3.5 | 83/100 |
| 19 | Gemini 3.1 Flash Lite | 80/100 |
“API cost” is what software developers pay to build on a model — ignore it if you just use the website. Each model answers each task once. Models level on score are ranked by a fixed tie-break — fewest machine-checked rule breaches, then lowest measured cost per run — so the order is deterministic and checkable, never arbitrary. Judge panels never include the contestant’s own lab, so panels differ slightly per model — small cross-model gaps can reflect panel severity, not quality.
Made by OpenAI. You use it inside ChatGPT — nothing to install.
Strongest showing: “A hundred words exactly” — scored 10/10 by the panel. Weakest: “Kill the purple” at 9/10. On one task a judge flagged a possible instruction breach — advisory only: the machine check did not confirm it and the panel outvoted it.
“The response perfectly follows all instructions, including the strict 100-word limit and the omission of a title. The story is engaging, well-written, and highly creative.”— google/gemini-3.1-pro-preview, judging blind · full receipts ↓
12 tasks · 16 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $5 in / $30 out per 1M tokens · full model page →
Made by OpenAI — their newest model. You use it inside ChatGPT — nothing to install.
Strongest showing: “The piano argument” — scored 10/10 by the panel. Weakest: “A hundred words exactly” at 8/10. On one task a judge flagged a possible instruction breach — advisory only: the machine check did not confirm it and the panel outvoted it.
“The response perfectly follows all instructions. It is exactly 14 lines of pure dialogue, reveals both siblings' valid points and misconceptions, and masterfully escalates the conflict to underlying family resentments.”— google/gemini-3.1-pro-preview, judging blind · full receipts ↓
12 tasks · 16 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $1 in / $6 out per 1M tokens · full model page →
Made by OpenAI — their newest model. You use it inside ChatGPT — nothing to install.
Strongest showing: “The piano argument” — scored 10/10 by the panel. Weakest: “A hundred words exactly” at 8/10. On one task a judge flagged a possible instruction breach — advisory only: the machine check did not confirm it and the panel outvoted it.
“The response perfectly follows all instructions, including the 12-16 line constraint (14 lines). It uses only dialogue, establishes each sibling's right/wrong points, and seamlessly shifts the conflict to their deeper familial resentments.”— google/gemini-3.1-pro-preview, judging blind · full receipts ↓
12 tasks · 16 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $5 in / $30 out per 1M tokens · full model page →
See it for yourself
One task from the suite, and what the leader, the middle of the table and the bottom actually wrote. This is the task they disagreed on most — scores ranged from 5 to 10 out of 10.
Write a story of exactly 100 words about a lost glove. No title — just the story. Count carefully: exactly 100.
At dawn, Mira found a red glove lying beneath the station clock, palm upward, as if asking for snow. She slipped it into her coat and boarded the train, meaning to leave it with lost property. But the glove warmed and curled around her fingers, tugging whenever the train passed certain fields. At th…
Judge: The response perfectly follows all instructions, including the strict 100-word limit and the omission of a title. The story is engaging, well-written, and highly creative.
The lost glove lay in the snow by the bus stop. A child had dropped it while boarding. Wind blew it under a bush. Days passed. A dog found it and carried it home. The owner tossed it aside. Rain soaked it. Then a woman walking her path spotted the bright blue fabric. She picked it up, cleaned it, an…
Judge: The response perfectly follows all instructions, including the exact 100-word count constraint and the omission of a title. The story is clear, well-written, and directly addresses the prompt.
The wool glove lay abandoned on the icy pavement, a lonely orphan of winter. It had once hugged a warm hand, feeling the pulse of life and the gentle squeeze of a secret held tight. Now, it rested in a puddle of slush, its navy fibers darkening with freezing water. A passerby hurried past, eyes fixe…
Judge: Word count is 101, not exactly 100 as required; otherwise decent narrative quality but fails the explicit constraint.
Trimmed to fit. Every full answer, every judge’s score: GPT-5.5, Grok 4.5, Gemini 3.1 Flash Lite
How this ranking is made
Every model answers the identical published task suite. Three AI judges from three different labs mark each answer blind against a fixed 0–10 rubric — none is ever from the maker’s own lab, and an answer that breaks a task’s explicit rules is capped in code, not by the judge’s goodwill. The score is the average, out of 100.
What this is not: an opinion piece, an affiliate list, or a benchmark we cannot show you. Every score links its raw outputs and judge verdicts. The full protocol · receipts: GPT-5.5, GPT-5.6 Terra, GPT-5.6 Sol, Gemini 3.5 Flash, Qwen3.7 Max, Gemini 3.1 Pro Preview, GPT-5.6 Luna, GPT-5.3-Codex, DeepSeek V4 Pro, Grok 4.5, Claude Sonnet 5, Claude Opus 4.8, Claude Fable 5, Kimi K3, Claude Opus 4.6, GLM 5.2, DeepSeek V4 Flash, Mistral Medium 3.5, Gemini 3.1 Flash Lite
Questions people ask
What is the best AI for creative writing in 2026?
GPT-5.5 leads our tested ranking with 98/100 on our 12-task creative-writing suite (12 tasks), in a dead heat with GPT-5.6 Terra (97). Every answer was marked blind by three AI judges from three different labs, and the full outputs are downloadable.
How is this ranking made?
Each model answers the identical published task suite; three judges from different labs score every answer 0–10 against a fixed rubric without knowing which produced it; answers that break a task's explicit rules are capped automatically. The score is the average, out of 100. No vendor pays for placement.
How often does this page update?
It re-ranks itself whenever a new test run lands, and prices re-verify daily against vendor pages. The current ranking was last computed on 16 Aug 2026.
Head-to-heads in creative writing
Show all 20 tested pairs ▾
More rankings ▾
Best AI for writing · Best AI chatbot for everyday use · Best AI for coding · Best free AI model · Best AI for spreadsheets and Excel · Best AI essay writer · Best AI for summarising documents · Best AI for extracting data from text · Best AI for writing emails · Best AI for everyday maths and percentages · Best AI for customer service replies · Best AI for revision and study notes · Best AI for vibe coding · Best AI for making flashcards · Best AI for social media posts · Best AI for job applications and cover letters · Best AI for presentations · Best AI for writing your CV · Best AI for research skills · Best AI for translation · Best AI for travel planning · Best AI for emotional support · Best AI for everyday legal questions · Best AI for health questions · Best AI for writing a book · Best AI for humanising AI text · Best-value AI model API · every model we track · every tool