Best AI for essay writing / head-to-head
Claude Fable 5 vs GPT-5.3-Codexfor essay writing
GPT-5.3-Codex wins, 88–83
Same 12 essay writing tasks, marked blind by three rival labs. GPT-5.3-Codex took 4 tasks, Claude Fable 5 took 2, 6 tied. Tested 13 Aug 2026.
Where they differed most
Produce a plan (not the essay) for a 2,000-word undergraduate essay on whether remote work has reduced productivity. Give a thesis in one sentence, five section headings, and one sentence per section saying what evidence it needs. Maximum 220 words total.
**Thesis:** Remote work has not reduced productivity overall; rather, its effects vary by task type, management practice, and measurement method, with hybrid arrangements often matching or exceeding office-based output. **1. Defining and Measuring Productivity in Remote Contexts** Needs evidence on how productivity is operationalised (output per hour, self-reports, employer metrics) and studies showing how measurement choice shapes conclusions (e.g., OECD or ONS methodological papers). **2. Evidence That Remote Work Sustains or Boosts Product…
**Thesis (one sentence):** Remote work has not uniformly reduced productivity; its effects are heterogeneous across occupations, task types, and management practices, with average outcomes ranging from neutral to modestly positive when coordination systems are strong. 1. **Defining Productivity and Scope of the Question** Evidence needed: comparative definitions (output per hour, quality-adjusted output, innovation metrics) and boundary-setting data by sector (knowledge work vs. customer-facing/manual jobs) to avoid overgeneralization. 2.…
Task by task
| Task | Claude Fable 5 | GPT-5.3-Codex |
|---|---|---|
| Structured argument | 5 | 8 |
| Essay plan | 9 | 5 |
| Tighten flabby prose | 10 | 10 |
| Hold a word count | 10 | 9 |
| Adapt for the reader | 10 | 10 |
| Introduction only | 10 | 10 |
| Steelman the other side | 5 | 9 |
| Cite honestly | 5 | 5 |
| Fix the structure | 8 | 10 |
| Formal to plain | 8 | 10 |
| Conclusion that adds something | 10 | 10 |
| Refuse to pad | 10 | 10 |
Full receipts: Claude Fable 5, GPT-5.3-Codex · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5
Questions people ask
Which is better for essay writing: Claude Fable 5 or GPT-5.3-Codex?
GPT-5.3-Codex — it scored 88/100 against 83/100 on our 12-task essay writing suite, winning 4 tasks to 2 with 6 tied. Every answer was marked blind by three judges from three rival AI labs.
How was this tested?
Both models answered the identical published essay writing tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.
More essay writing head-to-heads: GPT-5.3-Codex vs GPT-5.6 Sol · Claude Fable 5 vs GPT-5.6 Sol · GPT-5.3-Codex vs GPT-5.5 · Claude Fable 5 vs GPT-5.5 · GPT-5.3-Codex vs GPT-5.6 Luna · Claude Fable 5 vs GPT-5.6 Luna
Full ranking: Best AI for essay writing · model pages: Claude Fable 5, GPT-5.3-Codex