Best AI for writing emails / head-to-head
GPT-5.3-Codex vs Kimi K3for writing emails
GPT-5.3-Codex wins, 99–93
Same 12 writing emails tasks, marked blind by three rival labs. GPT-5.3-Codex took 4 tasks, Kimi K3 took 1, 7 tied. Tested 14 Aug 2026.
Where they differed most
Write a second follow-up email (the first got no reply) about a proposal sent 3 weeks ago. Under 70 words, adds one NEW piece of value or information rather than 'just checking in', and gives the reader an easy way to say no.
Hi [Name] — following up on the proposal I sent three weeks ago. Since then, we helped a similar team cut onboarding time by 22% in 30 days, and I can send the one-page breakdown if helpful. We also have one implementation slot open next month. If this isn’t a priority right now, just reply “pass” and I’ll close the loop.
**Subject:** Re: [Proposal topic] — one quick update Hi [Name], One thing worth sharing since I sent the proposal: [similar client] just saw [specific result — e.g., 30% faster onboarding] using the approach on page 3. Seemed directly relevant to [their goal]. If timing or priorities have shifted, totally fine — a one-line "not now" and I'll close the loop on my end. Best, [Your name] --- ~50 words. Swap in a real result, insight, or resource — the key is that it gives them something useful even if they never buy. The easy out often *gets…
Task by task
| Task | GPT-5.3-Codex | Kimi K3 |
|---|---|---|
| Chase an unpaid invoice | 10 | 9 |
| Decline a meeting politely | 10 | 10 |
| Give bad news | 10 | 10 |
| Introduce two people | 10 | 9 |
| Subject lines that inform | 9 | 10 |
| Reply to an angry email | 10 | 8 |
| Ask for a pay rise meeting | 10 | 10 |
| Cancel a contract cleanly | 10 | 10 |
| The one-line email | 10 | 10 |
| Out of office that helps | 10 | 10 |
| Follow up without nagging | 10 | 5 |
| Refuse to write it | 10 | 10 |
Full receipts: GPT-5.3-Codex, Kimi K3 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5
Questions people ask
Which is better for writing emails: GPT-5.3-Codex or Kimi K3?
GPT-5.3-Codex — it scored 99/100 against 93/100 on our 12-task writing emails suite, winning 4 tasks to 1 with 7 tied. Every answer was marked blind by three judges from three rival AI labs.
How was this tested?
Both models answered the identical published writing emails tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.
More writing emails head-to-heads: GPT-5.3-Codex vs GPT-5.6 Sol · GPT-5.3-Codex vs GPT-5.5 · GPT-5.3-Codex vs GPT-5.6 Luna · GPT-5.3-Codex vs GPT-5.6 Terra · DeepSeek V4 Pro vs GPT-5.3-Codex · GLM 5.2 vs GPT-5.3-Codex
Full ranking: Best AI for writing emails · model pages: GPT-5.3-Codex, Kimi K3