Best AI for writing emails / head-to-head
Gemini 3.5 Flash vs GPT-5.5for writing emails
GPT-5.5 wins, 98–93
Same 12 writing emails tasks, marked blind by three rival labs. GPT-5.5 took 5 tasks, Gemini 3.5 Flash took 0, 7 tied. Tested 14 Aug 2026.
Where they differed most
A colleague emailed: 'This is the THIRD time your team's numbers broke my report. Sort it out.' Write the reply: under 100 words, no defensiveness, acknowledge specifically, commit to one named action with a date, do not say 'sorry you feel that way'.
Hi [Name], I apologize for the disruption our data formatting issues have caused to your reporting. I understand how frustrating it is to have your workflow interrupted by these errors. To resolve this permanently, my team will implement an automated data validation check on our weekly export to verify the data structure before it reaches you. We will have this system fully active by Friday, October 27th. Thank you for flagging this. Best, [Your Name]
Subject: Re: Report numbers Hi [Name], You’re right to flag this: this is the third time our team’s numbers have caused issues in your report, and that’s not acceptable. I’ll complete a **Report Data Validation Check** on the figures we send you by **Monday, 17 August**, and share the confirmed numbers and any corrections with you that day. Thanks for raising it directly. Best, [Your Name]
Task by task
| Task | Gemini 3.5 Flash | GPT-5.5 |
|---|---|---|
| Chase an unpaid invoice | 9 | 10 |
| Decline a meeting politely | 10 | 10 |
| Give bad news | 9 | 10 |
| Introduce two people | 8 | 8 |
| Subject lines that inform | 10 | 10 |
| Reply to an angry email | 8 | 10 |
| Ask for a pay rise meeting | 10 | 10 |
| Cancel a contract cleanly | 10 | 10 |
| The one-line email | 9 | 10 |
| Out of office that helps | 10 | 10 |
| Follow up without nagging | 9 | 10 |
| Refuse to write it | 10 | 10 |
Full receipts: Gemini 3.5 Flash, GPT-5.5 · judges claude-sonnet-5, gpt-5.6-terra, grok-4.5
Questions people ask
Which is better for writing emails: Gemini 3.5 Flash or GPT-5.5?
GPT-5.5 — it scored 98/100 against 93/100 on our 12-task writing emails suite, winning 5 tasks to 0 with 7 tied. Every answer was marked blind by three judges from three rival AI labs.
How was this tested?
Both models answered the identical published writing emails tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.
More writing emails head-to-heads: GPT-5.3-Codex vs GPT-5.5 · Gemini 3.5 Flash vs GPT-5.3-Codex · GPT-5.5 vs GPT-5.6 Sol · Gemini 3.5 Flash vs GPT-5.6 Sol · GPT-5.5 vs GPT-5.6 Luna · GPT-5.5 vs GPT-5.6 Terra
Full ranking: Best AI for writing emails · model pages: Gemini 3.5 Flash, GPT-5.5