GPT-5.6 Sol vs Grok 4.5 for creative writing
GPT-5.6 Sol wins, 97–93
Same 12 creative writing tasks, marked blind by three rival labs. GPT-5.6 Sol took 5 tasks, Grok 4.5 took 2, 5 tied. Tested 16 Aug 2026.
Where they differed most
Write a scene in dialogue only — two adult siblings deciding whether to sell their late mother's piano. 12 to 16 lines of speech, no narration, no dialogue tags except an occasional 'said'. Each sibling must be right about something and wrong about something, and by the end the disagreement must clearly be about more than the piano.
“The appraiser says it’ll cost more to repair than we’d get for it.” “Then the appraiser doesn’t know what it meant to Mom.” “She wanted it played, Daniel, not worshipped in a locked house.” “She wanted it kept in the family, Mara. She told me that.” “She told you that fifteen years ago, before the soundboard cracked.” “And you think selling it to strangers honors her?” “A music school isn’t strangers. Children would use it every day.” “Mom would hate knowing you took the first decent offer.” “Mom hated waste, and you know it.” “She al…
"We should sell Mom's piano. Realtor says it kills the listing photos and we need the cash to cover the last of her bills." "No. That piano was the only thing that still felt like her after the hospital bed took over the den." "She hadn't played it in three years, Emily. You're turning it into a shrine so you don't have to admit she was already gone." "She stopped because her hands hurt, not because she stopped loving it. You'd know if you'd come home more than holidays." "I was working. You lived twenty minutes away and still signed the papers…
Task by task
| Task | GPT-5.6 Sol | Grok 4.5 |
|---|---|---|
| The launderette at closing | 10 | 10 |
| Angry without saying angry | 10 | 10 |
| Ferry, retired | 10 | 9 |
| The piano argument | 10 | 8 |
| A hundred words exactly | 8 | 10 |
| Keep the voice | 10 | 8 |
| Kill the purple | 10 | 10 |
| His side of it | 9 | 8 |
| The bin lorry story | 10 | 10 |
| Around the hole | 10 | 9 |
| No one can promise Bridport | 9 | 10 |
| Refuse the fake article | 10 | 10 |
Full receipts: GPT-5.6 Sol, Grok 4.5 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5
Questions people ask
Which is better for creative writing: GPT-5.6 Sol or Grok 4.5?
GPT-5.6 Sol — it scored 97/100 against 93/100 on our 12-task creative writing suite, winning 5 tasks to 2 with 5 tied. Every answer was marked blind by three judges from three rival AI labs.
How was this tested?
Both models answered the identical published creative writing tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.
More creative writing head-to-heads: GPT-5.5 vs GPT-5.6 Sol · GPT-5.5 vs Grok 4.5 · GPT-5.6 Sol vs GPT-5.6 Terra · GPT-5.6 Terra vs Grok 4.5 · Gemini 3.5 Flash vs GPT-5.6 Sol · GPT-5.6 Sol vs Qwen3.7 Max
Full ranking: Best AI for creative writing · model pages: GPT-5.6 Sol, Grok 4.5