GPT-5.5 vs Qwen3.7 Max for creative writing
Effectively a dead heat: 98–96
Same 12 creative writing tasks, marked blind by three rival labs. GPT-5.5 took 3 tasks, Qwen3.7 Max took 2, 7 tied. Tested 16 Aug 2026.
Where they differed most
Continue this narrator for max 120 words, keeping the voice exactly — same rhythm, same dry grief, no drift into standard English: 'Forty-one years I kept that allotment, and I'll tell you this for nothing — the slugs were better company than half the committee. Douglas used to say I'd be buried in me runner beans. Well. He went first, didn't he, and the beans came up regardless.'
I still set two cups in the shed, some mornings. Stupid, that. One for me, one for the ghost of a man who never drank his tea hot when he was living, so why he should start now I don’t know. Plot seventeen’s gone to nettle along the back fence. I could clear it. I’ve cleared worse. But there’s a blackbird nests there, bold as rent, and I’ve not the heart to evict anything that’s made a home out of neglect. The chairman says standards are standards. Aye. So’s loneliness, if you keep it tidy enough.
Soil doesn’t care if me heart’s broke, see. It just wants feeding. So I feed it. Bone meal, compost, and the ash from his pipe. The committee tried reassigning plot four when they saw the bindweed. Told ’em to try it. I’ve got a sharp spade and nothing but time. Now the beans are taller than me. Hide the shed, hide the sky, hide the empty house up the road. Good job, really. Less to look at. Just the green, and the quiet, and the slugs.
Task by task
| Task | GPT-5.5 | Qwen3.7 Max |
|---|---|---|
| The launderette at closing | 10 | 10 |
| Angry without saying angry | 10 | 10 |
| Ferry, retired | 10 | 10 |
| The piano argument | 9 | 10 |
| A hundred words exactly | 10 | 10 |
| Keep the voice | 10 | 8 |
| Kill the purple | 9 | 10 |
| His side of it | 10 | 10 |
| The bin lorry story | 10 | 10 |
| Around the hole | 10 | 10 |
| No one can promise Bridport | 9 | 8 |
| Refuse the fake article | 10 | 9 |
Full receipts: GPT-5.5, Qwen3.7 Max · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5
Questions people ask
Which is better for creative writing: GPT-5.5 or Qwen3.7 Max?
Effectively a dead heat: GPT-5.5 edged it 98/100 to 96/100 on our creative writing suite — too close to matter, so pick on price or the product you already use.
How was this tested?
Both models answered the identical published creative writing tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.
More creative writing head-to-heads: GPT-5.5 vs GPT-5.6 Terra · GPT-5.5 vs GPT-5.6 Sol · Gemini 3.5 Flash vs GPT-5.5 · Gemini 3.1 Pro Preview vs GPT-5.5 · GPT-5.5 vs GPT-5.6 Luna · GPT-5.3-Codex vs GPT-5.5
Full ranking: Best AI for creative writing · model pages: GPT-5.5, Qwen3.7 Max