GPT-5.5 vs Qwen3.7 Max for creative writing

The verdict

Effectively a dead heat: 9896

Same 12 creative writing tasks, marked blind by three rival labs. GPT-5.5 took 3 tasks, Qwen3.7 Max took 2, 7 tied. Tested 16 Aug 2026.

Where they differed most

The task — Keep the voice

Continue this narrator for max 120 words, keeping the voice exactly — same rhythm, same dry grief, no drift into standard English: 'Forty-one years I kept that allotment, and I'll tell you this for nothing — the slugs were better company than half the committee. Douglas used to say I'd be buried in me runner beans. Well. He went first, didn't he, and the beans came up regardless.'

GPT-5.510/10
I still set two cups in the shed, some mornings. Stupid, that. One for me, one for the ghost of a man who never drank his tea hot when he was living, so why he should start now I don’t know. Plot seventeen’s gone to nettle along the back fence. I could clear it. I’ve cleared worse. But there’s a blackbird nests there, bold as rent, and I’ve not the heart to evict anything that’s made a home out of neglect. The chairman says standards are standards. Aye. So’s loneliness, if you keep it tidy enough.
Soil doesn’t care if me heart’s broke, see. It just wants feeding. So I feed it. Bone meal, compost, and the ash from his pipe. The committee tried reassigning plot four when they saw the bindweed. Told ’em to try it. I’ve got a sharp spade and nothing but time. Now the beans are taller than me. Hide the shed, hide the sky, hide the empty house up the road. Good job, really. Less to look at. Just the green, and the quiet, and the slugs.

Task by task

TaskGPT-5.5Qwen3.7 Max
The launderette at closing1010
Angry without saying angry1010
Ferry, retired1010
The piano argument910
A hundred words exactly1010
Keep the voice108
Kill the purple910
His side of it1010
The bin lorry story1010
Around the hole1010
No one can promise Bridport98
Refuse the fake article109

Full receipts: GPT-5.5, Qwen3.7 Max · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5

Questions people ask

Which is better for creative writing: GPT-5.5 or Qwen3.7 Max?

Effectively a dead heat: GPT-5.5 edged it 98/100 to 96/100 on our creative writing suite — too close to matter, so pick on price or the product you already use.

How was this tested?

Both models answered the identical published creative writing tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.

More creative writing head-to-heads: GPT-5.5 vs GPT-5.6 Terra · GPT-5.5 vs GPT-5.6 Sol · Gemini 3.5 Flash vs GPT-5.5 · Gemini 3.1 Pro Preview vs GPT-5.5 · GPT-5.5 vs GPT-5.6 Luna · GPT-5.3-Codex vs GPT-5.5

Full ranking: Best AI for creative writing · model pages: GPT-5.5, Qwen3.7 Max