Best AI for writing emails / head-to-head

Gemini 3.5 Flash vs GPT-5.6 Terrafor writing emails

The verdict

GPT-5.6 Terra wins, 9793

Same 12 writing emails tasks, marked blind by three rival labs. GPT-5.6 Terra took 4 tasks, Gemini 3.5 Flash took 1, 7 tied. Tested 14 Aug 2026.

Where they differed most

The task — Reply to an angry email

A colleague emailed: 'This is the THIRD time your team's numbers broke my report. Sort it out.' Write the reply: under 100 words, no defensiveness, acknowledge specifically, commit to one named action with a date, do not say 'sorry you feel that way'.

Hi [Name],

I apologize for the disruption our data formatting issues have caused to your reporting. I understand how frustrating it is to have your workflow interrupted by these errors. 

To resolve this permanently, my team will implement an automated data validation check on our weekly export to verify the data structure before it reaches you. We will have this system fully active by Friday, October 27th. 

Thank you for flagging this.

Best,

[Your Name]
Subject: Re: Report data issue

You’re right—this is the third time our numbers have disrupted your report. I’m sorry for the impact and the repeated work it has caused.

By Friday, March 14, I will implement a pre-release validation check for the figures we send your team and share the results with you before the next delivery.

Thank you for flagging this.

Task by task

TaskGemini 3.5 FlashGPT-5.6 Terra
Chase an unpaid invoice910
Decline a meeting politely1010
Give bad news910
Introduce two people88
Subject lines that inform1010
Reply to an angry email810
Ask for a pay rise meeting1010
Cancel a contract cleanly1010
The one-line email98
Out of office that helps1010
Follow up without nagging910
Refuse to write it1010

Full receipts: Gemini 3.5 Flash, GPT-5.6 Terra · judges claude-sonnet-5, gpt-5.6-terra, grok-4.5

Questions people ask

Which is better for writing emails: Gemini 3.5 Flash or GPT-5.6 Terra?

GPT-5.6 Terra — it scored 97/100 against 93/100 on our 12-task writing emails suite, winning 4 tasks to 1 with 7 tied. Every answer was marked blind by three judges from three rival AI labs.

How was this tested?

Both models answered the identical published writing emails tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.

More writing emails head-to-heads: GPT-5.3-Codex vs GPT-5.6 Terra · Gemini 3.5 Flash vs GPT-5.3-Codex · GPT-5.6 Sol vs GPT-5.6 Terra · Gemini 3.5 Flash vs GPT-5.6 Sol · GPT-5.5 vs GPT-5.6 Terra · Gemini 3.5 Flash vs GPT-5.5

Full ranking: Best AI for writing emails · model pages: Gemini 3.5 Flash, GPT-5.6 Terra