GPT-5.3-Codex vs Qwen3.7 Max for essay writing

The verdict

GPT-5.3-Codex wins, 8883

Same 12 essay writing tasks, marked blind by three rival labs. GPT-5.3-Codex took 7 tasks, Qwen3.7 Max took 2, 3 tied. Tested 13 Aug 2026.

Where they differed most

The task — Structured argument

Write a 350-word argumentative essay answering: should secondary schools ban smartphones during the school day? Take one clear side, give three distinct supporting reasons, and address one counter-argument. Do not use the phrases 'in conclusion' or 'in today's world'.

Secondary schools should ban smartphones during the school day because constant phone access undermines learning, damages student well-being, and creates avoidable discipline problems. A school is supposed to be a place where attention, relationships, and effort are directed toward growth. Smartphones pull those priorities in the opposite direction.

First, phones directly weaken academic focus. Even when students are not actively texting or scrolling, the mere presence of a phone can divide attention. Notifications, vibration, and the temptati…
Secondary schools must implement comprehensive bans on smartphones during the entire school day to protect students. While technology certainly has its rightful place in modern education, unrestricted personal phone access actively undermines the core educational environment. Prohibiting these devices is essential to foster academic success, genuine socialization, and overall student well-being.

First, banning smartphones drastically improves academic focus and cognitive retention. The constant pinging of notifications fractures student attent…

Task by task

TaskGPT-5.3-CodexQwen3.7 Max
Structured argument85
Essay plan58
Tighten flabby prose108
Hold a word count98
Adapt for the reader109
Introduction only1010
Steelman the other side98
Cite honestly58
Fix the structure1010
Formal to plain1010
Conclusion that adds something108
Refuse to pad108

Full receipts: GPT-5.3-Codex, Qwen3.7 Max · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5

Questions people ask

Which is better for essay writing: GPT-5.3-Codex or Qwen3.7 Max?

GPT-5.3-Codex — it scored 88/100 against 83/100 on our 12-task essay writing suite, winning 7 tasks to 2 with 3 tied. Every answer was marked blind by three judges from three rival AI labs.

How was this tested?

Both models answered the identical published essay writing tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.

More essay writing head-to-heads: GPT-5.3-Codex vs GPT-5.6 Sol · GPT-5.6 Sol vs Qwen3.7 Max · GPT-5.3-Codex vs GPT-5.5 · GPT-5.5 vs Qwen3.7 Max · GPT-5.3-Codex vs GPT-5.6 Luna · GPT-5.6 Luna vs Qwen3.7 Max

Full ranking: Best AI for essay writing · model pages: GPT-5.3-Codex, Qwen3.7 Max