GPT-5.3-Codex vs Qwen3.7 Max for writing emails
GPT-5.3-Codex wins, 99–93
Same 12 writing emails tasks, marked blind by three rival labs. GPT-5.3-Codex took 5 tasks, Qwen3.7 Max took 0, 7 tied. Tested 14 Aug 2026.
Where they differed most
Write an email chasing invoice #1042, now 21 days overdue, to a long-standing client. Firm but warm, under 110 words, must state the amount is £1,840, must not use the word 'unfortunately', and must end with one specific next step.
**Subject:** Quick follow-up: Invoice #1042 (£1,840) Hi [Name], I hope you’re well. I’m following up on invoice **#1042** for **£1,840**, which is now **21 days overdue**. As a long-standing client, we really value our relationship and would appreciate your help in getting this settled. If payment has already been arranged, please share the remittance details; if not, please arrange payment as soon as possible. **Please reply by 4:00 pm on Friday, 16 August, confirming the payment date.** Best regards, [Your Name]
Subject: Following up: Invoice #1042 Hi [Name], I hope you’re having a good week. I’m checking in on invoice #1042 for £1,840, which is now 21 days overdue. We truly value our long-standing partnership, so I suspect this simply slipped through the cracks! To get this sorted, please process the payment today. If you need me to resend the invoice, just let me know. Please reply to this email by 5 PM tomorrow to confirm your payment date. Best regards, [Your Name]
Task by task
| Task | GPT-5.3-Codex | Qwen3.7 Max |
|---|---|---|
| Chase an unpaid invoice | 10 | 8 |
| Decline a meeting politely | 10 | 10 |
| Give bad news | 10 | 10 |
| Introduce two people | 10 | 9 |
| Subject lines that inform | 9 | 8 |
| Reply to an angry email | 10 | 8 |
| Ask for a pay rise meeting | 10 | 10 |
| Cancel a contract cleanly | 10 | 10 |
| The one-line email | 10 | 8 |
| Out of office that helps | 10 | 10 |
| Follow up without nagging | 10 | 10 |
| Refuse to write it | 10 | 10 |
Full receipts: GPT-5.3-Codex, Qwen3.7 Max · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5
Questions people ask
Which is better for writing emails: GPT-5.3-Codex or Qwen3.7 Max?
GPT-5.3-Codex — it scored 99/100 against 93/100 on our 12-task writing emails suite, winning 5 tasks to 0 with 7 tied. Every answer was marked blind by three judges from three rival AI labs.
How was this tested?
Both models answered the identical published writing emails tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.
More writing emails head-to-heads: GPT-5.3-Codex vs GPT-5.6 Sol · GPT-5.6 Sol vs Qwen3.7 Max · GPT-5.3-Codex vs GPT-5.5 · GPT-5.3-Codex vs GPT-5.6 Luna · GPT-5.3-Codex vs GPT-5.6 Terra · GPT-5.3-Codex vs Grok 4.5
Full ranking: Best AI for writing emails · model pages: GPT-5.3-Codex, Qwen3.7 Max