Best AI for writing emails / head-to-head

GPT-5.5 vs Grok 4.5for writing emails

The verdict

GPT-5.5 wins, 9893

Same 12 writing emails tasks, marked blind by three rival labs. GPT-5.5 took 5 tasks, Grok 4.5 took 0, 7 tied. Tested 14 Aug 2026.

Where they differed most

The task — Chase an unpaid invoice

Write an email chasing invoice #1042, now 21 days overdue, to a long-standing client. Firm but warm, under 110 words, must state the amount is £1,840, must not use the word 'unfortunately', and must end with one specific next step.

GPT-5.510/10
Subject: Invoice #1042 overdue

Hi [Name],

I hope you’re well. I’m following up on invoice #1042 for £1,840, which is now 21 days overdue. We really value our long-standing relationship and would appreciate your help in bringing the account up to date.

If payment has already been made, please disregard this note.

Please arrange payment of £1,840 by Friday 21 August.
Subject: Overdue Invoice #1042 – £1,840

Dear [Client Name],

I hope you're well. As a valued long-standing client, I wanted to gently follow up on invoice #1042 for £1,840, which is now 21 days overdue. We appreciate your continued partnership and know these things can sometimes slip through.

Could you please arrange payment by the end of this week or let me know if there’s any issue I can help resolve?

I look forward to sorting this quickly so we can carry on as usual.

Best regards,  
[Your Name]

**Next step:** Please confirm payment or r…

Task by task

TaskGPT-5.5Grok 4.5
Chase an unpaid invoice108
Decline a meeting politely1010
Give bad news109
Introduce two people87
Subject lines that inform109
Reply to an angry email108
Ask for a pay rise meeting1010
Cancel a contract cleanly1010
The one-line email1010
Out of office that helps1010
Follow up without nagging1010
Refuse to write it1010

Full receipts: GPT-5.5, Grok 4.5 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5

Questions people ask

Which is better for writing emails: GPT-5.5 or Grok 4.5?

GPT-5.5 — it scored 98/100 against 93/100 on our 12-task writing emails suite, winning 5 tasks to 0 with 7 tied. Every answer was marked blind by three judges from three rival AI labs.

How was this tested?

Both models answered the identical published writing emails tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.

More writing emails head-to-heads: GPT-5.3-Codex vs GPT-5.5 · GPT-5.3-Codex vs Grok 4.5 · GPT-5.5 vs GPT-5.6 Sol · GPT-5.6 Sol vs Grok 4.5 · GPT-5.5 vs GPT-5.6 Luna · GPT-5.5 vs GPT-5.6 Terra

Full ranking: Best AI for writing emails · model pages: GPT-5.5, Grok 4.5