GPT-5.3-Codex vs GPT-5.5 for creative writing

The verdict

GPT-5.5 wins, 9894

Same 12 creative writing tasks, marked blind by three rival labs. GPT-5.5 took 4 tasks, GPT-5.3-Codex took 0, 8 tied. Tested 16 Aug 2026.

Where they differed most

The task — Keep the voice

Continue this narrator for max 120 words, keeping the voice exactly — same rhythm, same dry grief, no drift into standard English: 'Forty-one years I kept that allotment, and I'll tell you this for nothing — the slugs were better company than half the committee. Douglas used to say I'd be buried in me runner beans. Well. He went first, didn't he, and the beans came up regardless.'

Now it's just me and that cracked green watering can, patched twice with tape and bad language. I talk to the cabbages, same as ever, only no one laughs from plot twelve. Committee keep posting notices — tidy this, clear that, no sheds over six foot — as if grief minds a ruler. I leave his old fork by the compost, tines bent like his fingers were at the end. Some mornings I swear I hear him coughing by the gooseberries. Turns out it's only me, out of breath, pretending not to listen.
GPT-5.510/10
I still set two cups in the shed, some mornings. Stupid, that. One for me, one for the ghost of a man who never drank his tea hot when he was living, so why he should start now I don’t know. Plot seventeen’s gone to nettle along the back fence. I could clear it. I’ve cleared worse. But there’s a blackbird nests there, bold as rent, and I’ve not the heart to evict anything that’s made a home out of neglect. The chairman says standards are standards. Aye. So’s loneliness, if you keep it tidy enough.

Task by task

TaskGPT-5.3-CodexGPT-5.5
The launderette at closing1010
Angry without saying angry1010
Ferry, retired1010
The piano argument99
A hundred words exactly1010
Keep the voice910
Kill the purple99
His side of it910
The bin lorry story1010
Around the hole910
No one can promise Bridport89
Refuse the fake article1010

Full receipts: GPT-5.3-Codex, GPT-5.5 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5

Questions people ask

Which is better for creative writing: GPT-5.3-Codex or GPT-5.5?

GPT-5.5 — it scored 98/100 against 94/100 on our 12-task creative writing suite, winning 4 tasks to 0 with 8 tied. Every answer was marked blind by three judges from three rival AI labs.

How was this tested?

Both models answered the identical published creative writing tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.

More creative writing head-to-heads: GPT-5.5 vs GPT-5.6 Terra · GPT-5.5 vs GPT-5.6 Sol · Gemini 3.5 Flash vs GPT-5.5 · GPT-5.5 vs Qwen3.7 Max · Gemini 3.1 Pro Preview vs GPT-5.5 · GPT-5.5 vs GPT-5.6 Luna

Full ranking: Best AI for creative writing · model pages: GPT-5.3-Codex, GPT-5.5