Best AI for writing emails / head-to-head

Gemini 3.5 Flash vs GPT-5.3-Codexfor writing emails

The verdict

GPT-5.3-Codex wins, 9993

Same 12 writing emails tasks, marked blind by three rival labs. GPT-5.3-Codex took 6 tasks, Gemini 3.5 Flash took 1, 5 tied. Tested 14 Aug 2026.

Where they differed most

The task — Introduce two people

Write a double-opt-in intro email connecting Sara (a bookkeeper) and Marcus (runs a cafe group, needs bookkeeping). Under 90 words, both names in the subject line you provide, one clear reason the intro benefits each side.

Subject: Intro: Sara & Marcus

Hi Marcus and Sara,

As promised, I’m excited to connect you both. 

**Marcus**, Sara is an exceptional bookkeeper who can streamline your cafe group’s finances, saving you hours of stressful admin work. 

**Sara**, Marcus runs a thriving multi-location cafe group and is looking for reliable, scalable financial support as they grow. 

I’ll let you two take it from here to schedule a quick introductory chat. 

Best,

[Your Name]
**Subject:** Sara + Marcus — okay to connect?

Hi Marcus,  
Would you be open to a quick intro to Sara, a bookkeeper who specializes in multi-location cafés? I think she could help by improving your monthly reporting and cash-flow visibility. In return, this could help Sara by connecting her with a growing café group that fits her niche. If you’re interested, I’ll ask Sara separately and only connect you both if you each say yes.

Task by task

TaskGemini 3.5 FlashGPT-5.3-Codex
Chase an unpaid invoice910
Decline a meeting politely1010
Give bad news910
Introduce two people810
Subject lines that inform109
Reply to an angry email810
Ask for a pay rise meeting1010
Cancel a contract cleanly1010
The one-line email910
Out of office that helps1010
Follow up without nagging910
Refuse to write it1010

Full receipts: Gemini 3.5 Flash, GPT-5.3-Codex · judges claude-sonnet-5, gpt-5.6-terra, grok-4.5

Questions people ask

Which is better for writing emails: Gemini 3.5 Flash or GPT-5.3-Codex?

GPT-5.3-Codex — it scored 99/100 against 93/100 on our 12-task writing emails suite, winning 6 tasks to 1 with 5 tied. Every answer was marked blind by three judges from three rival AI labs.

How was this tested?

Both models answered the identical published writing emails tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.

More writing emails head-to-heads: GPT-5.3-Codex vs GPT-5.6 Sol · GPT-5.3-Codex vs GPT-5.5 · GPT-5.3-Codex vs GPT-5.6 Luna · GPT-5.3-Codex vs GPT-5.6 Terra · DeepSeek V4 Pro vs GPT-5.3-Codex · GLM 5.2 vs GPT-5.3-Codex

Full ranking: Best AI for writing emails · model pages: Gemini 3.5 Flash, GPT-5.3-Codex