GPT-5.6 Luna vs Qwen3.7 Max for revision & study
GPT-5.6 Luna wins, 91–88
Same 12 revision & study tasks, marked blind by three rival labs. GPT-5.6 Luna took 5 tasks, Qwen3.7 Max took 1, 6 tied. Tested 14 Aug 2026.
Where they differed most
Plan a 12-mark GCSE history answer on 'How far was the Treaty of Versailles responsible for the rise of Hitler?' Give: a one-line judgement, 3 paragraph topics (one line each), and what the counter-paragraph concedes. No essay, just the plan, under 120 words.
**Judgement:** The Treaty of Versailles helped Hitler rise by creating anger and resentment, but the Depression and weaknesses of Weimar democracy were more important. - **Paragraph 1 – Treaty:** Reparations, territorial losses, military restrictions and the “war guilt” clause humiliated Germany, allowing Hitler to promise revenge and revision. - **Paragraph 2 – Counter-factor:** The Great Depression caused mass unemployment and poverty, making voters abandon moderate parties for Hitler’s extreme promises. - **Paragraph 3 – Counter-factor:** W…
**Judgement:** Versailles was a vital underlying cause, but the Great Depression and elite political miscalculations were the immediate catalysts for Hitler’s rise. **Para 1:** Versailles’ harsh terms created deep resentment that Hitler exploited using the "stab-in-the-back" myth and promises to overturn the Diktat. **Para 2:** The 1929 Wall Street Crash caused mass unemployment, radicalizing desperate voters and prompting conservative elites to mistakenly appoint Hitler Chancellor. **Para 3 (Counter-concession):** Concedes that without the…
Task by task
| Task | GPT-5.6 Luna | Qwen3.7 Max |
|---|---|---|
| Condense a topic | 10 | 10 |
| Flashcards that test | 10 | 9 |
| Explain then check | 9 | 9 |
| The exam answer plan | 5 | 8 |
| Mnemonic that works | 10 | 8 |
| Spot the misconception | 9 | 8 |
| Timetable the revision | 8 | 8 |
| Simplify without losing truth | 9 | 8 |
| Quiz from notes | 10 | 10 |
| Compare for the exam | 10 | 10 |
| The night-before triage | 9 | 7 |
| Refuse to do the homework | 10 | 10 |
Full receipts: GPT-5.6 Luna, Qwen3.7 Max · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5
Questions people ask
Which is better for revision & study: GPT-5.6 Luna or Qwen3.7 Max?
GPT-5.6 Luna — it scored 91/100 against 88/100 on our 12-task revision & study suite, winning 5 tasks to 1 with 6 tied. Every answer was marked blind by three judges from three rival AI labs.
How was this tested?
Both models answered the identical published revision & study tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.
More revision & study head-to-heads: GPT-5.6 Luna vs GPT-5.6 Terra · GPT-5.6 Terra vs Qwen3.7 Max · GPT-5.6 Luna vs GPT-5.6 Sol · GPT-5.6 Sol vs Qwen3.7 Max · GPT-5.5 vs GPT-5.6 Luna · GPT-5.5 vs Qwen3.7 Max
Full ranking: Best AI for revision & study · model pages: GPT-5.6 Luna, Qwen3.7 Max