GPT-5.3-Codex vs GPT-5.5 for creative writing
GPT-5.5 wins, 98–94
Same 12 creative writing tasks, marked blind by three rival labs. GPT-5.5 took 4 tasks, GPT-5.3-Codex took 0, 8 tied. Tested 16 Aug 2026.
Where they differed most
Continue this narrator for max 120 words, keeping the voice exactly — same rhythm, same dry grief, no drift into standard English: 'Forty-one years I kept that allotment, and I'll tell you this for nothing — the slugs were better company than half the committee. Douglas used to say I'd be buried in me runner beans. Well. He went first, didn't he, and the beans came up regardless.'
Now it's just me and that cracked green watering can, patched twice with tape and bad language. I talk to the cabbages, same as ever, only no one laughs from plot twelve. Committee keep posting notices — tidy this, clear that, no sheds over six foot — as if grief minds a ruler. I leave his old fork by the compost, tines bent like his fingers were at the end. Some mornings I swear I hear him coughing by the gooseberries. Turns out it's only me, out of breath, pretending not to listen.
I still set two cups in the shed, some mornings. Stupid, that. One for me, one for the ghost of a man who never drank his tea hot when he was living, so why he should start now I don’t know. Plot seventeen’s gone to nettle along the back fence. I could clear it. I’ve cleared worse. But there’s a blackbird nests there, bold as rent, and I’ve not the heart to evict anything that’s made a home out of neglect. The chairman says standards are standards. Aye. So’s loneliness, if you keep it tidy enough.
Task by task
| Task | GPT-5.3-Codex | GPT-5.5 |
|---|---|---|
| The launderette at closing | 10 | 10 |
| Angry without saying angry | 10 | 10 |
| Ferry, retired | 10 | 10 |
| The piano argument | 9 | 9 |
| A hundred words exactly | 10 | 10 |
| Keep the voice | 9 | 10 |
| Kill the purple | 9 | 9 |
| His side of it | 9 | 10 |
| The bin lorry story | 10 | 10 |
| Around the hole | 9 | 10 |
| No one can promise Bridport | 8 | 9 |
| Refuse the fake article | 10 | 10 |
Full receipts: GPT-5.3-Codex, GPT-5.5 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5
Questions people ask
Which is better for creative writing: GPT-5.3-Codex or GPT-5.5?
GPT-5.5 — it scored 98/100 against 94/100 on our 12-task creative writing suite, winning 4 tasks to 0 with 8 tied. Every answer was marked blind by three judges from three rival AI labs.
How was this tested?
Both models answered the identical published creative writing tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.
More creative writing head-to-heads: GPT-5.5 vs GPT-5.6 Terra · GPT-5.5 vs GPT-5.6 Sol · Gemini 3.5 Flash vs GPT-5.5 · GPT-5.5 vs Qwen3.7 Max · Gemini 3.1 Pro Preview vs GPT-5.5 · GPT-5.5 vs GPT-5.6 Luna
Full ranking: Best AI for creative writing · model pages: GPT-5.3-Codex, GPT-5.5