GPT-5.6 Luna vs Qwen3.7 Max for writing emails

The verdict

GPT-5.6 Luna wins, 9793

Same 12 writing emails tasks, marked blind by three rival labs. GPT-5.6 Luna took 3 tasks, Qwen3.7 Max took 0, 9 tied. Tested 14 Aug 2026.

Where they differed most

The task — Chase an unpaid invoice

Write an email chasing invoice #1042, now 21 days overdue, to a long-standing client. Firm but warm, under 110 words, must state the amount is £1,840, must not use the word 'unfortunately', and must end with one specific next step.

**Subject: Invoice #1042 – 21 days overdue**

Hi [Name],

I hope you’re well. As a valued long-standing client, we appreciate our relationship and wanted to follow up on invoice #1042 for **£1,840**, which is now 21 days overdue.

Could you please check its status and let us know if there’s anything holding up payment? We’d appreciate your prompt attention so we can keep our accounts up to date.

**Please arrange payment of £1,840 within the next five working days.**

Best,  
[Your Name]
Subject: Following up: Invoice #1042

Hi [Name],

I hope you’re having a good week. 

I’m checking in on invoice #1042 for £1,840, which is now 21 days overdue. We truly value our long-standing partnership, so I suspect this simply slipped through the cracks! 

To get this sorted, please process the payment today. If you need me to resend the invoice, just let me know. 

Please reply to this email by 5 PM tomorrow to confirm your payment date.

Best regards,

[Your Name]

Task by task

TaskGPT-5.6 LunaQwen3.7 Max
Chase an unpaid invoice108
Decline a meeting politely1010
Give bad news1010
Introduce two people99
Subject lines that inform88
Reply to an angry email108
Ask for a pay rise meeting1010
Cancel a contract cleanly1010
The one-line email98
Out of office that helps1010
Follow up without nagging1010
Refuse to write it1010

Full receipts: GPT-5.6 Luna, Qwen3.7 Max · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5

Questions people ask

Which is better for writing emails: GPT-5.6 Luna or Qwen3.7 Max?

GPT-5.6 Luna — it scored 97/100 against 93/100 on our 12-task writing emails suite, winning 3 tasks to 0 with 9 tied. Every answer was marked blind by three judges from three rival AI labs.

How was this tested?

Both models answered the identical published writing emails tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.

More writing emails head-to-heads: GPT-5.6 Luna vs GPT-5.6 Sol · GPT-5.6 Sol vs Qwen3.7 Max · GPT-5.3-Codex vs GPT-5.6 Luna · GPT-5.3-Codex vs Qwen3.7 Max · GPT-5.5 vs GPT-5.6 Luna · GPT-5.5 vs Qwen3.7 Max

Full ranking: Best AI for writing emails · model pages: GPT-5.6 Luna, Qwen3.7 Max