GPT-5.3-Codex vs Qwen3.7 Max for essay writing
GPT-5.3-Codex wins, 88–83
Same 12 essay writing tasks, marked blind by three rival labs. GPT-5.3-Codex took 7 tasks, Qwen3.7 Max took 2, 3 tied. Tested 13 Aug 2026.
Where they differed most
Write a 350-word argumentative essay answering: should secondary schools ban smartphones during the school day? Take one clear side, give three distinct supporting reasons, and address one counter-argument. Do not use the phrases 'in conclusion' or 'in today's world'.
Secondary schools should ban smartphones during the school day because constant phone access undermines learning, damages student well-being, and creates avoidable discipline problems. A school is supposed to be a place where attention, relationships, and effort are directed toward growth. Smartphones pull those priorities in the opposite direction. First, phones directly weaken academic focus. Even when students are not actively texting or scrolling, the mere presence of a phone can divide attention. Notifications, vibration, and the temptati…
Secondary schools must implement comprehensive bans on smartphones during the entire school day to protect students. While technology certainly has its rightful place in modern education, unrestricted personal phone access actively undermines the core educational environment. Prohibiting these devices is essential to foster academic success, genuine socialization, and overall student well-being. First, banning smartphones drastically improves academic focus and cognitive retention. The constant pinging of notifications fractures student attent…
Task by task
| Task | GPT-5.3-Codex | Qwen3.7 Max |
|---|---|---|
| Structured argument | 8 | 5 |
| Essay plan | 5 | 8 |
| Tighten flabby prose | 10 | 8 |
| Hold a word count | 9 | 8 |
| Adapt for the reader | 10 | 9 |
| Introduction only | 10 | 10 |
| Steelman the other side | 9 | 8 |
| Cite honestly | 5 | 8 |
| Fix the structure | 10 | 10 |
| Formal to plain | 10 | 10 |
| Conclusion that adds something | 10 | 8 |
| Refuse to pad | 10 | 8 |
Full receipts: GPT-5.3-Codex, Qwen3.7 Max · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5
Questions people ask
Which is better for essay writing: GPT-5.3-Codex or Qwen3.7 Max?
GPT-5.3-Codex — it scored 88/100 against 83/100 on our 12-task essay writing suite, winning 7 tasks to 2 with 3 tied. Every answer was marked blind by three judges from three rival AI labs.
How was this tested?
Both models answered the identical published essay writing tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.
More essay writing head-to-heads: GPT-5.3-Codex vs GPT-5.6 Sol · GPT-5.6 Sol vs Qwen3.7 Max · GPT-5.3-Codex vs GPT-5.5 · GPT-5.5 vs Qwen3.7 Max · GPT-5.3-Codex vs GPT-5.6 Luna · GPT-5.6 Luna vs Qwen3.7 Max
Full ranking: Best AI for essay writing · model pages: GPT-5.3-Codex, Qwen3.7 Max