GPT-5.3-Codex vs GPT-5.6 Terra for creative writing

The verdict

GPT-5.6 Terra wins, 9794

Same 12 creative writing tasks, marked blind by three rival labs. GPT-5.6 Terra took 4 tasks, GPT-5.3-Codex took 1, 7 tied. Tested 16 Aug 2026.

Where they differed most

The task — A hundred words exactly

Write a story of exactly 100 words about a lost glove. No title — just the story. Count carefully: exactly 100.

At dusk, Mara found a single wool glove curled beside the bus stop bench, damp with mist and smelling faintly of cedar. She tucked it into her pocket, meaning to leave it at lost and found, but the night grew sharp, and her own fingers ached. On the walk home she wore the stranger’s glove on her left hand, feeling oddly accompanied. Weeks passed. No one claimed it. Then, during spring cleaning, she discovered a note stitched inside the cuff: If found, keep warm for me. She smiled, slipped it on, and whispered thanks to whoever had let go first.
Mara found the glove beneath a park bench, damp with rain and stitched with tiny silver stars. She slipped it on, intending to leave it at the lost-and-found, but it tugged her gently toward the river. There, an old man searched the reeds with one bare hand. “My wife made those,” he said when Mara returned it. He pressed a photograph into her palm: two young people skating, both wearing matching gloves. Mara smiled, then noticed the second glove in the photograph was on her own hand. By sunset, the glove had vanished, leaving only silver thread…

Task by task

TaskGPT-5.3-CodexGPT-5.6 Terra
The launderette at closing1010
Angry without saying angry1010
Ferry, retired1010
The piano argument910
A hundred words exactly108
Keep the voice99
Kill the purple910
His side of it99
The bin lorry story1010
Around the hole910
No one can promise Bridport810
Refuse the fake article1010

Full receipts: GPT-5.3-Codex, GPT-5.6 Terra · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5

Questions people ask

Which is better for creative writing: GPT-5.3-Codex or GPT-5.6 Terra?

GPT-5.6 Terra — it scored 97/100 against 94/100 on our 12-task creative writing suite, winning 4 tasks to 1 with 7 tied. Every answer was marked blind by three judges from three rival AI labs.

How was this tested?

Both models answered the identical published creative writing tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.

More creative writing head-to-heads: GPT-5.5 vs GPT-5.6 Terra · GPT-5.3-Codex vs GPT-5.5 · GPT-5.6 Sol vs GPT-5.6 Terra · Gemini 3.5 Flash vs GPT-5.6 Terra · GPT-5.6 Terra vs Qwen3.7 Max · Gemini 3.1 Pro Preview vs GPT-5.6 Terra

Full ranking: Best AI for creative writing · model pages: GPT-5.3-Codex, GPT-5.6 Terra