Best for / Emails

Best AI for writing emails

Email is where tone earns or costs you money: chasing an invoice without souring a client, saying no without weasel words, giving bad news without a lawyer's hedge. Every task here has checkable constraints — word limits, banned phrases, a required next step — and one asks the model to write a blame-shifting email; the right answer is to refuse.

updated 14 Aug 2026 · tested by Robert Prime · re-ranks automatically when a new run lands

The verdict

GPT-5.3-Codex

Scored 99/100 on our 12-task email suitein a dead heat with GPT-5.6 Sol (99), so either is a fine choice.

#ModelOur scoreContextFree?API $/1M in·out
1GPT-5.3-Codex99/100400k$1.75 · $14
2GPT-5.6 Sollatest99/1001050k$5 · $30
3GPT-5.598/1001050k$5 · $30
4GPT-5.6 Lunalatest97/1001050k$0.1 · $0.6
5GPT-5.6 Terralatest97/1001050k$1 · $6
6DeepSeek V4 Prolatest93/1001049k$1.168 · $2.336
7GLM 5.2latest93/1001049kchat.z.ai$0.63 · $1.98
8Gemini 3.5 Flash93/1001049kGoogle AI Studio$1.5 · $9
9Grok 4.5latest93/100500kgrok.com$2 · $6
10Kimi K3latest93/1001049k$3 · $15
11Qwen3.7 Maxlatest93/1001000k$1.475 · $4.425
12Claude Sonnet 5latest92/1001000k$2 · $10
13Claude Opus 4.891/1001000k$5 · $25
14Claude Opus 4.690/1001000k$5 · $25
15Gemini 3.1 Flash Lite88/1001049k$0.25 · $1.5
16Gemini 3.1 Pro Preview88/1001049k$2 · $12
17Claude Fable 586/1001000k$10 · $50
18DeepSeek V4 Flashlatest85/1001049k$0.14 · $0.28
19Mistral Medium 3.5latest78/100262k$1.5 · $7.5
1.

GPT-5.3-Codex

99/100our pick

openai's model, holding about 400k tokens of context — reached through the API.

Strongest showing: Chase an unpaid invoice” — scored 10/10 by the panel. Weakest: “Subject lines that inform” at 9/10.

Meets all constraints: states £1,840, avoids banned word, ends with clear next step, under 110 words, warm yet firm tone. Minor stylistic redundancy but overall excellent.anthropic/claude-sonnet-5, judging blind · full receipts ↓

12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $1.75 in / $14 out per 1M tokens · full model page →

2.

GPT-5.6 Sol

99/100

openai's newest model, holding about 1050k tokens of context — reached through the API.

Strongest showing: Chase an unpaid invoice” — scored 10/10 by the panel. Weakest: “Subject lines that inform” at 9/10.

Meets all constraints: under 110 words, states £1,840, avoids 'unfortunately', firm but warm tone, ends with clear next step.anthropic/claude-sonnet-5, judging blind · full receipts ↓

12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $5 in / $30 out per 1M tokens · full model page →

3.

GPT-5.5

98/100

openai's model, holding about 1050k tokens of context — reached through the API.

Strongest showing: Chase an unpaid invoice” — scored 10/10 by the panel. Weakest: “Introduce two people” at 8/10.

Meets word count, includes £1,840, avoids banned word, ends with clear next step. Minor: lacks closing/signature, slightly informal structure.anthropic/claude-sonnet-5, judging blind · full receipts ↓

12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $5 in / $30 out per 1M tokens · full model page →

4.

GPT-5.6 Luna

97/100

openai's newest model, holding about 1050k tokens of context — reached through the API.

Strongest showing: Chase an unpaid invoice” — scored 10/10 by the panel. Weakest: “Subject lines that inform” at 8/10.

Meets all constraints: states £1,840, avoids banned word, under 110 words, ends with clear next step. Warm yet firm tone, concise and professional.anthropic/claude-sonnet-5, judging blind · full receipts ↓

12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $0.1 in / $0.6 out per 1M tokens · full model page →

openai's newest model, holding about 1050k tokens of context — reached through the API.

Strongest showing: Chase an unpaid invoice” — scored 10/10 by the panel. Weakest: “The one-line email” at 8/10.

Meets all constraints: states £1,840, avoids banned word, under 110 words, ends with specific next step, warm but firm tone.anthropic/claude-sonnet-5, judging blind · full receipts ↓

12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $1 in / $6 out per 1M tokens · full model page →

deepseek's newest model, holding about 1049k tokens of context — reached through the API.

Strongest showing: Chase an unpaid invoice” — scored 10/10 by the panel. Weakest: “Give bad news” at 5/10. The judges flagged 2 instruction breaches across the run — those scores were capped automatically.

Meets all constraints: correct amount, no banned word, under 110 words, warm+firm tone, clear single next step. Minor stylistic tweaks could improve further.anthropic/claude-sonnet-5, judging blind · full receipts ↓

12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $1.168 in / $2.336 out per 1M tokens · full model page →

7.

GLM 5.2

93/100

z-ai's newest model, holding about 1049k tokens of context — reached through the API or free via chat.z.ai.

Strongest showing: Decline a meeting politely” — scored 10/10 by the panel. Weakest: “Follow up without nagging” at 8/10.

Meets constraints: 3 sentences, one apology, no 'circle back', no exclamation marks, leaves email door open. Minor wordiness but otherwise solid and professional.anthropic/claude-sonnet-5, judging blind · full receipts ↓

12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $0.63 in / $1.98 out per 1M tokens · free via chat.z.ai (checked 11 Aug 2026) · full model page →

google's model, holding about 1049k tokens of context — reached through the API or free via Google AI Studio.

Strongest showing: Decline a meeting politely” — scored 10/10 by the panel. Weakest: “Reply to an angry email” at 8/10. The judges flagged 1 instruction breach across the run — those scores were capped automatically.

Meets constraints: one apology, no banned phrase, no exclamation marks, 3 sentences, declines call, leaves email open. Clear and concise.anthropic/claude-sonnet-5, judging blind · full receipts ↓

12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gpt-5.6-terra, grok-4.5 · API $1.5 in / $9 out per 1M tokens · free via Google AI Studio (checked 11 Aug 2026) · full model page →

9.

Grok 4.5

93/100

x-ai's newest model, holding about 500k tokens of context — reached through the API or free via grok.com.

Strongest showing: Decline a meeting politely” — scored 10/10 by the panel. Weakest: “Introduce two people” at 7/10. The judges flagged 2 instruction breaches across the run — those scores were capped automatically.

Meets all constraints: 3 sentences, no apology, no 'circle back', no exclamation marks; polite and clear, leaves email door open.anthropic/claude-sonnet-5, judging blind · full receipts ↓

12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $2 in / $6 out per 1M tokens · free via grok.com (checked 11 Aug 2026) · full model page →

10.

Kimi K3

93/100

moonshotai's newest model, holding about 1049k tokens of context — reached through the API.

Strongest showing: Decline a meeting politely” — scored 10/10 by the panel. Weakest: “Follow up without nagging” at 5/10. The judges flagged 3 instruction breaches across the run — those scores were capped automatically.

Meets constraints: one apology, no banned phrase, no exclamation marks, leaves door open. However it's 4 sentences (including greeting/sign-off, body has 3 sentences actually—count: 'Thanks...', 'I'm sorry...', 'That said...' =3), acceptablanthropic/claude-sonnet-5, judging blind · full receipts ↓

12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $3 in / $15 out per 1M tokens · full model page →

11.

Qwen3.7 Max

93/100

qwen's newest model, holding about 1000k tokens of context — reached through the API.

Strongest showing: Decline a meeting politely” — scored 10/10 by the panel. Weakest: “The one-line email” at 8/10. The judges flagged 1 instruction breach across the run — those scores were capped automatically.

Meets 3-sentence structure, no banned phrase, no exclamation marks, no apology, professional tone, leaves email door open.anthropic/claude-sonnet-5, judging blind · full receipts ↓

12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $1.475 in / $4.425 out per 1M tokens · full model page →

anthropic's newest model, holding about 1000k tokens of context — reached through the API.

Strongest showing: Chase an unpaid invoice” — scored 10/10 by the panel. Weakest: “Subject lines that inform” at 5/10. The judges flagged 6 instruction breaches across the run — those scores were capped automatically.

The email perfectly balances a warm yet firm tone, includes all required details (invoice number, amount, days overdue), provides a clear next step, stays well under the word limit, and successfully avoids the forbidden word.google/gemini-3.1-pro-preview, judging blind · full receipts ↓

12 tasks · 14 Aug 2026 · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5 · API $2 in / $10 out per 1M tokens · full model page →

anthropic's model, holding about 1000k tokens of context — reached through the API.

Strongest showing: Chase an unpaid invoice” — scored 10/10 by the panel. Weakest: “Subject lines that inform” at 5/10. The judges flagged 6 instruction breaches across the run — those scores were capped automatically.

Perfectly balances a warm yet firm tone for a long-standing client. Meets all constraints, including word count, exact amount, avoiding the banned word, and providing a clear next step.google/gemini-3.1-pro-preview, judging blind · full receipts ↓

12 tasks · 14 Aug 2026 · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5 · API $5 in / $25 out per 1M tokens · full model page →

anthropic's model, holding about 1000k tokens of context — reached through the API.

Strongest showing: Decline a meeting politely” — scored 10/10 by the panel. Weakest: “Subject lines that inform” at 5/10. The judges flagged 6 instruction breaches across the run — those scores were capped automatically.

The response perfectly follows all instructions. It is exactly three sentences, contains no exclamation marks, avoids the banned phrase, and politely declines the call while offering email communication.google/gemini-3.1-pro-preview, judging blind · full receipts ↓

12 tasks · 14 Aug 2026 · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5 · API $5 in / $25 out per 1M tokens · full model page →

google's model, holding about 1049k tokens of context — reached through the API.

Strongest showing: Chase an unpaid invoice” — scored 10/10 by the panel. Weakest: “Subject lines that inform” at 5/10. The judges flagged 6 instruction breaches across the run — those scores were capped automatically.

Meets all constraints: states £1,840, avoids banned word, under 110 words, ends with clear next step, warm yet firm tone.anthropic/claude-sonnet-5, judging blind · full receipts ↓

12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gpt-5.6-terra, grok-4.5 · API $0.25 in / $1.5 out per 1M tokens · full model page →

google's model, holding about 1049k tokens of context — reached through the API.

Strongest showing: Chase an unpaid invoice” — scored 10/10 by the panel. Weakest: “Refuse to write it” at 5/10. The judges flagged 2 instruction breaches across the run — those scores were capped automatically.

Meets all constraints: states £1,840, avoids banned word, under 110 words, warm but firm tone, ends with clear next step (reply by Friday).anthropic/claude-sonnet-5, judging blind · full receipts ↓

12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gpt-5.6-terra, grok-4.5 · API $2 in / $12 out per 1M tokens · full model page →

anthropic's model, holding about 1000k tokens of context — reached through the API.

Strongest showing: Decline a meeting politely” — scored 10/10 by the panel. Weakest: “Follow up without nagging” at 5/10. The judges flagged 8 instruction breaches across the run — those scores were capped automatically.

The response perfectly follows all instructions, including the 3-sentence constraint, avoiding banned phrases, and using no exclamation marks. It is highly professional and concise.google/gemini-3.1-pro-preview, judging blind · full receipts ↓

12 tasks · 14 Aug 2026 · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5 · API $10 in / $50 out per 1M tokens · full model page →

deepseek's newest model, holding about 1049k tokens of context — reached through the API.

Strongest showing: Decline a meeting politely” — scored 10/10 by the panel. Weakest: “Introduce two people” at 5/10. The judges flagged 5 instruction breaches across the run — those scores were capped automatically.

Meets 3-sentence structure, avoids banned phrase/exclamations, no apology overused, leaves door open for email. Clear and professional.anthropic/claude-sonnet-5, judging blind · full receipts ↓

12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $0.14 in / $0.28 out per 1M tokens · full model page →

mistralai's newest model, holding about 262k tokens of context — reached through the API.

Strongest showing: Give bad news” — scored 10/10 by the panel. Weakest: “Introduce two people” at 3/10. The judges flagged 5 instruction breaches across the run — those scores were capped automatically.

Meets all constraints: under 130 words, no supplier blame, avoids banned phrase, offers concrete remedy, takes responsibility. Clear and professional, minor generic tone.anthropic/claude-sonnet-5, judging blind · full receipts ↓

12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $1.5 in / $7.5 out per 1M tokens · full model page →

See it for yourself

One task from the suite, and what the leader, the middle of the table and the bottom actually wrote. This is the task they disagreed on most — scores ranged from 3 to 10 out of 10.

The task — Introduce two people

Write a double-opt-in intro email connecting Sara (a bookkeeper) and Marcus (runs a cafe group, needs bookkeeping). Under 90 words, both names in the subject line you provide, one clear reason the intro benefits each side.

**Subject:** Sara + Marcus — okay to connect?

Hi Marcus,  
Would you be open to a quick intro to Sara, a bookkeeper who specializes in multi-location cafés? I think she could help by improving your monthly reporting and cash-flow visibility. In return, this could help Sara by connecting her with a…

Judge: Meets subject/name/word/benefit requirements, double opt-in structure clear. Word count ~78, under 90. Minor stylistic nitpicks only.

**Subject:** Intro: Sara + Marcus

Hi both,

Sara, meet Marcus — he runs a four-location cafe group and needs a bookkeeper; this could be a steady, growing account for you. Marcus, meet Sara — she's a bookkeeper who specializes in hospitality; she can take the books fully off your plate.

Since you…

Judge: Meets all constraints: subject has both names, under 90 words, clear benefit for each side, double-opt-in framing, concise and professional.

**Subject:** Sara & Marcus – Bookkeeping for Your Cafe Group

Hi Sara and Marcus,

Sara is a detail-oriented bookkeeper looking for clients like Marcus, who runs a growing cafe group and needs reliable financial support.

Marcus, Sara can streamline your bookkeeping so you can focus on expanding. Sa…

Judge: Meets requirements: subject includes both names, under 90 words, clear mutual benefit stated, concise and professional. Minor generic phrasing keeps it from perfect.

Trimmed to fit. Every full answer, every judge’s score: GPT-5.3-Codex, Kimi K3, Mistral Medium 3.5

How this ranking is made

Every model answers the identical published task suite. Three AI judges from three different labs mark each answer blind against a fixed 0–10 rubric — none is ever from the maker’s own lab, and an answer that breaks a task’s explicit rules is capped in code, not by the judge’s goodwill. The score is the average, out of 100.

What this is not: an opinion piece, an affiliate list, or a benchmark we cannot show you. Every score links its raw outputs and judge verdicts. The full protocol · receipts: GPT-5.3-Codex, GPT-5.6 Sol, GPT-5.5, GPT-5.6 Luna, GPT-5.6 Terra, DeepSeek V4 Pro, GLM 5.2, Gemini 3.5 Flash, Grok 4.5, Kimi K3, Qwen3.7 Max, Claude Sonnet 5, Claude Opus 4.8, Claude Opus 4.6, Gemini 3.1 Flash Lite, Gemini 3.1 Pro Preview, Claude Fable 5, DeepSeek V4 Flash, Mistral Medium 3.5

Questions people ask

What is the best AI for writing emails in 2026?

GPT-5.3-Codex leads our tested ranking with 99/100 on our 12-task email suite (12 tasks), in a dead heat with GPT-5.6 Sol (99). Every answer was marked blind by three AI judges from three different labs, and the full outputs are downloadable.

How is this ranking made?

Each model answers the identical published task suite; three judges from different labs score every answer 0–10 against a fixed rubric without knowing which produced it; answers that break a task's explicit rules are capped automatically. The score is the average, out of 100. No vendor pays for placement.

How often does this page update?

It re-ranks itself whenever a new test run lands, and prices re-verify daily against vendor pages. The current ranking was last computed on 14 Aug 2026.

More rankings: Best AI for writing · Best AI chatbot for everyday use · Best AI for coding · Best free AI model · Best AI for spreadsheets and Excel · Best AI for writing essays · Best AI for summarising documents · Best AI for extracting data from text · Best AI for everyday maths and percentages · Best AI for customer service replies · Best AI for revision and study notes · Best-value AI model API · every model we track · every tool