GPT-5.6 Luna vs GPT-5.6 Sol for creative writing

The verdict

GPT-5.6 Sol wins, 9794

Same 12 creative writing tasks, marked blind by three rival labs. GPT-5.6 Sol took 4 tasks, GPT-5.6 Luna took 1, 7 tied. Tested 16 Aug 2026.

Where they differed most

The task — A hundred words exactly

Write a story of exactly 100 words about a lost glove. No title — just the story. Count carefully: exactly 100.

At dawn, Mara found one red glove beneath the station bench. Its twin had vanished during yesterday’s snowstorm, leaving her hand cold and her heart oddly colder. She tucked the survivor into her pocket, intending to mend the pair someday. On the train, a little boy began to cry. He had lost his glove, blue with yellow stars. Mara offered hers. The boy smiled, but his mother shook her head. “It isn’t yours,” she said. “It was mine,” Mara replied, “but perhaps it has been looking for him.” Outside, snow covered the tracks, carrying secrets towar…
At dawn, Mara found a single red glove beside the frozen pond. Its fingers pointed toward the reeds, as if asking her to follow. She did, stepping over silver grass, until she heard a faint whimper. Beneath an overturned boat crouched a shivering boy, one hand bare, the other wearing the glove’s twin. Mara wrapped him in her coat and led him home. Years later, the boy returned, carrying that red glove in a glass case. “You saved me,” he said. Mara smiled. “No,” she replied, “your glove did.” Outside, snow began falling, soft as feathers, coveri…

Task by task

TaskGPT-5.6 LunaGPT-5.6 Sol
The launderette at closing910
Angry without saying angry1010
Ferry, retired1010
The piano argument910
A hundred words exactly108
Keep the voice910
Kill the purple810
His side of it99
The bin lorry story1010
Around the hole1010
No one can promise Bridport99
Refuse the fake article1010

Full receipts: GPT-5.6 Luna, GPT-5.6 Sol · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5

Questions people ask

Which is better for creative writing: GPT-5.6 Luna or GPT-5.6 Sol?

GPT-5.6 Sol — it scored 97/100 against 94/100 on our 12-task creative writing suite, winning 4 tasks to 1 with 7 tied. Every answer was marked blind by three judges from three rival AI labs.

How was this tested?

Both models answered the identical published creative writing tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.

More creative writing head-to-heads: GPT-5.5 vs GPT-5.6 Sol · GPT-5.5 vs GPT-5.6 Luna · GPT-5.6 Sol vs GPT-5.6 Terra · GPT-5.6 Luna vs GPT-5.6 Terra · Gemini 3.5 Flash vs GPT-5.6 Sol · GPT-5.6 Sol vs Qwen3.7 Max

Full ranking: Best AI for creative writing · model pages: GPT-5.6 Luna, GPT-5.6 Sol