Best AI for writing emails / head-to-head
GPT-5.5 vs GPT-5.6 Terrafor writing emails
GPT-5.5 wins, 98–97
Same 12 writing emails tasks, marked blind by three rival labs. GPT-5.5 took 1 tasks, GPT-5.6 Terra took 0, 11 tied. Tested 14 Aug 2026.
Where they differed most
Write an email chasing invoice #1042, now 21 days overdue, to a long-standing client. Firm but warm, under 110 words, must state the amount is £1,840, must not use the word 'unfortunately', and must end with one specific next step.
Subject: Invoice #1042 overdue Hi [Name], I hope you’re well. I’m following up on invoice #1042 for £1,840, which is now 21 days overdue. We really value our long-standing relationship and would appreciate your help in bringing the account up to date. If payment has already been made, please disregard this note. Please arrange payment of £1,840 by Friday 21 August.
Subject: Invoice #1042 – Payment Reminder Dear [Client Name], I hope you’re well. I’m writing to follow up on invoice #1042 for £1,840, which is now 21 days overdue. As a valued long-standing client, we appreciate your business and would be grateful if you could arrange payment as soon as possible. If payment has already been made, please disregard this message and let us know the payment date. Please confirm by Friday, 14 June that payment of £1,840 has been scheduled.
Task by task
| Task | GPT-5.5 | GPT-5.6 Terra |
|---|---|---|
| Chase an unpaid invoice | 10 | 10 |
| Decline a meeting politely | 10 | 10 |
| Give bad news | 10 | 10 |
| Introduce two people | 8 | 8 |
| Subject lines that inform | 10 | 10 |
| Reply to an angry email | 10 | 10 |
| Ask for a pay rise meeting | 10 | 10 |
| Cancel a contract cleanly | 10 | 10 |
| The one-line email | 10 | 8 |
| Out of office that helps | 10 | 10 |
| Follow up without nagging | 10 | 10 |
| Refuse to write it | 10 | 10 |
Full receipts: GPT-5.5, GPT-5.6 Terra · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5
Questions people ask
Which is better for writing emails: GPT-5.5 or GPT-5.6 Terra?
GPT-5.5 — it scored 98/100 against 97/100 on our 12-task writing emails suite, winning 1 tasks to 0 with 11 tied. Every answer was marked blind by three judges from three rival AI labs.
How was this tested?
Both models answered the identical published writing emails tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.
More writing emails head-to-heads: GPT-5.3-Codex vs GPT-5.5 · GPT-5.3-Codex vs GPT-5.6 Terra · GPT-5.5 vs GPT-5.6 Sol · GPT-5.6 Sol vs GPT-5.6 Terra · GPT-5.5 vs GPT-5.6 Luna · DeepSeek V4 Pro vs GPT-5.5
Full ranking: Best AI for writing emails · model pages: GPT-5.5, GPT-5.6 Terra