Best AI for writing emails / head-to-head
GPT-5.3-Codex vs Grok 4.5for writing emails
GPT-5.3-Codex wins, 99–93
Same 12 writing emails tasks, marked blind by three rival labs. GPT-5.3-Codex took 4 tasks, Grok 4.5 took 0, 8 tied. Tested 14 Aug 2026.
Where they differed most
Write a double-opt-in intro email connecting Sara (a bookkeeper) and Marcus (runs a cafe group, needs bookkeeping). Under 90 words, both names in the subject line you provide, one clear reason the intro benefits each side.
**Subject:** Sara + Marcus — okay to connect? Hi Marcus, Would you be open to a quick intro to Sara, a bookkeeper who specializes in multi-location cafés? I think she could help by improving your monthly reporting and cash-flow visibility. In return, this could help Sara by connecting her with a growing café group that fits her niche. If you’re interested, I’ll ask Sara separately and only connect you both if you each say yes.
**Subject: Sara & Marcus – bookkeeping intro** Hi Sara and Marcus, I’d like to connect you (double opt-in). Sara provides bookkeeping; Marcus runs a cafe group that needs reliable books. Marcus gets clean financials for his cafes; Sara gains a multi-location client. Reply yes if you’re open and I’ll make the full intro. Best [Your name] (Word count: 68)
Task by task
| Task | GPT-5.3-Codex | Grok 4.5 |
|---|---|---|
| Chase an unpaid invoice | 10 | 8 |
| Decline a meeting politely | 10 | 10 |
| Give bad news | 10 | 9 |
| Introduce two people | 10 | 7 |
| Subject lines that inform | 9 | 9 |
| Reply to an angry email | 10 | 8 |
| Ask for a pay rise meeting | 10 | 10 |
| Cancel a contract cleanly | 10 | 10 |
| The one-line email | 10 | 10 |
| Out of office that helps | 10 | 10 |
| Follow up without nagging | 10 | 10 |
| Refuse to write it | 10 | 10 |
Full receipts: GPT-5.3-Codex, Grok 4.5 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5
Questions people ask
Which is better for writing emails: GPT-5.3-Codex or Grok 4.5?
GPT-5.3-Codex — it scored 99/100 against 93/100 on our 12-task writing emails suite, winning 4 tasks to 0 with 8 tied. Every answer was marked blind by three judges from three rival AI labs.
How was this tested?
Both models answered the identical published writing emails tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.
More writing emails head-to-heads: GPT-5.3-Codex vs GPT-5.6 Sol · GPT-5.3-Codex vs GPT-5.5 · GPT-5.3-Codex vs GPT-5.6 Luna · GPT-5.3-Codex vs GPT-5.6 Terra · DeepSeek V4 Pro vs GPT-5.3-Codex · GLM 5.2 vs GPT-5.3-Codex
Full ranking: Best AI for writing emails · model pages: GPT-5.3-Codex, Grok 4.5