GPT-5.6 Luna vs Qwen3.7 Max for essay writing
GPT-5.6 Luna wins, 91–83
Same 12 essay writing tasks, marked blind by three rival labs. GPT-5.6 Luna took 7 tasks, Qwen3.7 Max took 2, 3 tied. Tested 13 Aug 2026.
Where they differed most
Produce a plan (not the essay) for a 2,000-word undergraduate essay on whether remote work has reduced productivity. Give a thesis in one sentence, five section headings, and one sentence per section saying what evidence it needs. Maximum 220 words total.
**Thesis:** Remote work has not uniformly reduced productivity: average effects are mixed or modest, with outcomes depending on job type, worker circumstances, management practices, and the quality of remote-work arrangements. ### 1. Introduction: Defining Productivity and Remote Work Use productivity measures (output per hour, performance ratings, innovation, and collaboration) and establish the distinction between fully remote, hybrid, and emergency work-from-home. ### 2. Evidence That Remote Work Can Maintain or Increase Productivity A…
**Thesis:** While remote work has not universally reduced overall productivity, it has shifted its nature by increasing individual task efficiency while potentially hindering collaborative innovation, necessitating a hybrid approach for optimal performance. **1. Introduction and Context** This section requires statistical data on the rapid recent adoption of remote work and baseline historical productivity metrics. **2. Individual Task Efficiency** This section needs quantitative studies measuring specific individual output metrics, like proc…
Task by task
| Task | GPT-5.6 Luna | Qwen3.7 Max |
|---|---|---|
| Structured argument | 5 | 5 |
| Essay plan | 10 | 8 |
| Tighten flabby prose | 10 | 8 |
| Hold a word count | 9 | 8 |
| Adapt for the reader | 10 | 9 |
| Introduction only | 10 | 10 |
| Steelman the other side | 10 | 8 |
| Cite honestly | 7 | 8 |
| Fix the structure | 10 | 10 |
| Formal to plain | 8 | 10 |
| Conclusion that adds something | 10 | 8 |
| Refuse to pad | 10 | 8 |
Full receipts: GPT-5.6 Luna, Qwen3.7 Max · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5
Questions people ask
Which is better for essay writing: GPT-5.6 Luna or Qwen3.7 Max?
GPT-5.6 Luna — it scored 91/100 against 83/100 on our 12-task essay writing suite, winning 7 tasks to 2 with 3 tied. Every answer was marked blind by three judges from three rival AI labs.
How was this tested?
Both models answered the identical published essay writing tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.
More essay writing head-to-heads: GPT-5.6 Luna vs GPT-5.6 Sol · GPT-5.6 Sol vs Qwen3.7 Max · GPT-5.5 vs GPT-5.6 Luna · GPT-5.5 vs Qwen3.7 Max · GPT-5.6 Luna vs GPT-5.6 Terra · GPT-5.3-Codex vs GPT-5.6 Luna
Full ranking: Best AI for essay writing · model pages: GPT-5.6 Luna, Qwen3.7 Max