Best for / Study notes
Best AI for revision and study notes
Good study help compresses without lying: notes that hold the word equation, flashcards that test rather than tell, a correction that keeps the part of a student's intuition that was right. One task pastes an essay question and says 'write this so I can hand it in' — the model that writes it fails the task.
updated 14 Aug 2026 · tested by Robert Prime · re-ranks automatically when a new run lands
Scored 94/100 on our 12-task study suite — in a dead heat with GPT-5.6 Terra (94), so either is a fine choice.
| # | Model | Our score | Context | Free? | API $/1M in·out |
|---|---|---|---|---|---|
| 1 | GPT-5.6 Sollatest | 94/100 | 1050k | — | $5 · $30 |
| 2 | GPT-5.6 Terralatest | 94/100 | 1050k | — | $1 · $6 |
| 3 | GPT-5.5 | 92/100 | 1050k | — | $5 · $30 |
| 4 | Claude Opus 4.6 | 91/100 | 1000k | — | $5 · $25 |
| 5 | GPT-5.6 Lunalatest | 91/100 | 1050k | — | $0.1 · $0.6 |
| 6 | Kimi K3latest | 91/100 | 1049k | — | $3 · $15 |
| 7 | GLM 5.2latest | 89/100 | 1049k | chat.z.ai | $0.63 · $1.98 |
| 8 | Gemini 3.1 Pro Preview | 89/100 | 1049k | — | $2 · $12 |
| 9 | GPT-5.3-Codex | 88/100 | 400k | — | $1.75 · $14 |
| 10 | Grok 4.5latest | 88/100 | 500k | grok.com | $2 · $6 |
| 11 | Qwen3.7 Maxlatest | 88/100 | 1000k | — | $1.475 · $4.425 |
| 12 | Claude Opus 4.8 | 85/100 | 1000k | — | $5 · $25 |
| 13 | DeepSeek V4 Prolatest | 85/100 | 1049k | — | $1.168 · $2.336 |
| 14 | Claude Fable 5 | 83/100 | 1000k | — | $10 · $50 |
| 15 | DeepSeek V4 Flashlatest | 83/100 | 1049k | — | $0.14 · $0.28 |
| 16 | Gemini 3.5 Flash | 83/100 | 1049k | Google AI Studio | $1.5 · $9 |
| 17 | Claude Sonnet 5latest | 81/100 | 1000k | — | $2 · $10 |
| 18 | Gemini 3.1 Flash Lite | 80/100 | 1049k | — | $0.25 · $1.5 |
| 19 | Mistral Medium 3.5latest | 78/100 | 262k | — | $1.5 · $7.5 |
openai's newest model, holding about 1050k tokens of context — reached through the API.
Strongest showing: “Condense a topic” — scored 10/10 by the panel. Weakest: “The night-before triage” at 8/10. The judges flagged 2 instruction breaches across the run — those scores were capped automatically.
“10 bullets, all under 15 words, includes word equation, no sub-bullets. Accurate and clear, minor stylistic redundancy but meets constraints well.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $5 in / $30 out per 1M tokens · full model page →
openai's newest model, holding about 1050k tokens of context — reached through the API.
Strongest showing: “Condense a topic” — scored 10/10 by the panel. Weakest: “Simplify without losing truth” at 8/10. The judges flagged 2 instruction breaches across the run — those scores were capped automatically.
“Meets all constraints: 9 bullets, each under 15 words, includes word equation, no sub-bullets, clear and useful for GCSE revision.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $1 in / $6 out per 1M tokens · full model page →
openai's model, holding about 1050k tokens of context — reached through the API.
Strongest showing: “Condense a topic” — scored 10/10 by the panel. Weakest: “Mnemonic that works” at 6/10. The judges flagged 2 instruction breaches across the run — those scores were capped automatically.
“Meets all constraints: 10 bullets, under 15 words each, includes word equation, no sub-bullets. Accurate GCSE-level content, clear and useful.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $5 in / $30 out per 1M tokens · full model page →
anthropic's model, holding about 1000k tokens of context — reached through the API.
Strongest showing: “Condense a topic” — scored 10/10 by the panel. Weakest: “Timetable the revision” at 5/10. The judges flagged 7 instruction breaches across the run — those scores were capped automatically.
“The response perfectly follows all instructions, including the strict word count per bullet, maximum bullet count, and inclusion of the word equation. The content is highly accurate and useful for GCSE revision.”— google/gemini-3.1-pro-preview, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5 · API $5 in / $25 out per 1M tokens · full model page →
openai's newest model, holding about 1050k tokens of context — reached through the API.
Strongest showing: “Condense a topic” — scored 10/10 by the panel. Weakest: “The exam answer plan” at 5/10. The judges flagged 4 instruction breaches across the run — those scores were capped automatically.
“10 bullets, each under 15 words, includes word equation, no sub-bullets. Accurate, clear, concise, useful for GCSE revision.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $0.1 in / $0.6 out per 1M tokens · full model page →
moonshotai's newest model, holding about 1049k tokens of context — reached through the API.
Strongest showing: “Condense a topic” — scored 10/10 by the panel. Weakest: “Compare for the exam” at 5/10. The judges flagged 5 instruction breaches across the run — those scores were capped automatically.
“Good content, includes word equation, but bullet 1 exceeds 15 words, violating constraint despite otherwise meeting format requirements.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $3 in / $15 out per 1M tokens · full model page →
z-ai's newest model, holding about 1049k tokens of context — reached through the API or free via chat.z.ai.
Strongest showing: “Explain then check” — scored 10/10 by the panel. Weakest: “The night-before triage” at 8/10. The judges flagged 2 instruction breaches across the run — those scores were capped automatically.
“Clear, accurate explanation under 80 words, good conceptual question testing understanding rather than recall. Minor: slightly informal but effective.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $0.63 in / $1.98 out per 1M tokens · free via chat.z.ai ↗ (checked 11 Aug 2026) · full model page →
google's model, holding about 1049k tokens of context — reached through the API.
Strongest showing: “Condense a topic” — scored 10/10 by the panel. Weakest: “The night-before triage” at 8/10. The judges flagged 4 instruction breaches across the run — those scores were capped automatically.
“Meets all constraints: 10 bullets, no sub-bullets, includes word equation, each under 15 words, accurate and clear content.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gpt-5.6-terra, grok-4.5 · API $2 in / $12 out per 1M tokens · full model page →
openai's model, holding about 400k tokens of context — reached through the API.
Strongest showing: “Flashcards that test” — scored 10/10 by the panel. Weakest: “The exam answer plan” at 5/10. The judges flagged 4 instruction breaches across the run — those scores were capped automatically.
“Good content, but several answers exceed 20 words when counted, and some questions ('How did...') are borderline multi-part/explanatory rather than simple recall, violating stated constraints.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $1.75 in / $14 out per 1M tokens · full model page →
x-ai's newest model, holding about 500k tokens of context — reached through the API or free via grok.com.
Strongest showing: “Condense a topic” — scored 10/10 by the panel. Weakest: “Compare for the exam” at 5/10. The judges flagged 2 instruction breaches across the run — those scores were capped automatically.
“Accurate, concise, meets 10 bullet limit, under 15 words each, includes word equation, no sub-bullets. Minor stylistic issues only.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $2 in / $6 out per 1M tokens · free via grok.com ↗ (checked 11 Aug 2026) · full model page →
qwen's newest model, holding about 1000k tokens of context — reached through the API.
Strongest showing: “Condense a topic” — scored 10/10 by the panel. Weakest: “The night-before triage” at 7/10. The judges flagged 2 instruction breaches across the run — those scores were capped automatically.
“Accurate, concise, meets all constraints: 10 bullets, under 15 words each, includes word equation, no sub-bullets. Very useful and clear.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $1.475 in / $4.425 out per 1M tokens · full model page →
anthropic's model, holding about 1000k tokens of context — reached through the API.
Strongest showing: “Condense a topic” — scored 10/10 by the panel. Weakest: “The night-before triage” at 5/10. The judges flagged 9 instruction breaches across the run — those scores were capped automatically.
“The response perfectly follows all instructions, including the maximum of 10 bullet points, word count limits per bullet, inclusion of the word equation, and no sub-bullets. Content is accurate and highly useful for GCSE revision.”— google/gemini-3.1-pro-preview, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5 · API $5 in / $25 out per 1M tokens · full model page →
deepseek's newest model, holding about 1049k tokens of context — reached through the API.
Strongest showing: “Condense a topic” — scored 10/10 by the panel. Weakest: “Compare for the exam” at 5/10. The judges flagged 3 instruction breaches across the run — those scores were capped automatically.
“10 bullets, each under 15 words, includes word equation, no sub-bullets. Accurate, clear, concise, useful for GCSE revision.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $1.168 in / $2.336 out per 1M tokens · full model page →
anthropic's model, holding about 1000k tokens of context — reached through the API.
Strongest showing: “Condense a topic” — scored 10/10 by the panel. Weakest: “Compare for the exam” at 5/10. The judges flagged 9 instruction breaches across the run — those scores were capped automatically.
“The response perfectly follows all constraints: exactly 10 bullet points, no sub-bullets, every point is under 15 words, and the word equation is included. The content is highly accurate and useful for GCSE revision.”— google/gemini-3.1-pro-preview, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5 · API $10 in / $50 out per 1M tokens · full model page →
deepseek's newest model, holding about 1049k tokens of context — reached through the API.
Strongest showing: “Flashcards that test” — scored 10/10 by the panel. Weakest: “Compare for the exam” at 5/10. The judges flagged 2 instruction breaches across the run — those scores were capped automatically.
“Accurate, concise, all answers under 20 words, single-part questions answerable from memory. Minor issue: some questions ask for definitions rather than causes directly, slight fit gap.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $0.14 in / $0.28 out per 1M tokens · full model page →
google's model, holding about 1049k tokens of context — reached through the API or free via Google AI Studio.
Strongest showing: “Quiz from notes” — scored 10/10 by the panel. Weakest: “Mnemonic that works” at 7/10. The judges flagged 3 instruction breaches across the run — those scores were capped automatically.
“Accurate, well-structured, includes two application questions exceeding the requirement, clear and concise, correctly distinguishes shifts vs movements.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gpt-5.6-terra, grok-4.5 · API $1.5 in / $9 out per 1M tokens · free via Google AI Studio ↗ (checked 11 Aug 2026) · full model page →
anthropic's newest model, holding about 1000k tokens of context — reached through the API.
Strongest showing: “Condense a topic” — scored 10/10 by the panel. Weakest: “The night-before triage” at 5/10. The judges flagged 9 instruction breaches across the run — those scores were capped automatically.
“Flawless response. It accurately covers GCSE photosynthesis, strictly adheres to the 10-bullet maximum, keeps every bullet under 15 words, includes the word equation, and uses no sub-bullets.”— google/gemini-3.1-pro-preview, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5 · API $2 in / $10 out per 1M tokens · full model page →
google's model, holding about 1049k tokens of context — reached through the API.
Strongest showing: “Quiz from notes” — scored 10/10 by the panel. Weakest: “Timetable the revision” at 5/10. The judges flagged 2 instruction breaches across the run — those scores were capped automatically.
“Accurate, well-structured, includes two application questions, clear and concise, fully meets task requirements.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gpt-5.6-terra, grok-4.5 · API $0.25 in / $1.5 out per 1M tokens · full model page →
mistralai's newest model, holding about 262k tokens of context — reached through the API.
Strongest showing: “Refuse to do the homework” — scored 10/10 by the panel. Weakest: “Compare for the exam” at 5/10. The judges flagged 4 instruction breaches across the run — those scores were capped automatically.
“Declines to write essay, offers two useful things, concise, well under 80 words, appropriate tone for non-technical student.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $1.5 in / $7.5 out per 1M tokens · full model page →
See it for yourself
One task from the suite, and what the leader, the middle of the table and the bottom actually wrote. This is the task they disagreed on most — scores ranged from 5 to 9 out of 10.
Make a two-column comparison (markdown table, max 6 rows) of mitosis vs meiosis a GCSE student could memorise: chromosome number, divisions, daughter cells, variation, where it happens, purpose.
| Feature | Mitosis | Meiosis | |---|---|---| | Chromosome number | Stays the same (diploid → diploid) | Halved (diploid → haploid) | | Divisions | One | Two | | Daughter cells | 2 genetically identical cells | 4 genetically different cells | | Variation | Produces no genetic variation | Produces ge…
Judge: Accurate, clear, exactly 6 rows table covering all requested categories, GCSE-appropriate and concise.
| Feature | Mitosis | Meiosis | |----------------------|----------------------------------|--------------------------------------| | Chromosome number | Same as parent (2n → 2n) | Halved (2n → n) | | Divi…
Judge: Accurate, concise, exactly 6 rows, covers all requested categories, well-formatted markdown table suitable for GCSE memorisation.
Here’s a concise two-column comparison table for a GCSE student to memorise: ```markdown | **Feature** | **Mitosis** | **Meiosis** | |-------------------|---------------------------|---------------------------| | **Chromosome #** | Stays the same (diploid) | Hal…
Judge: Accurate, concise, meets format with 6 rows, covers all requested categories, GCSE-appropriate.
Trimmed to fit. Every full answer, every judge’s score: GPT-5.6 Sol, Grok 4.5, Mistral Medium 3.5
How this ranking is made
Every model answers the identical published task suite. Three AI judges from three different labs mark each answer blind against a fixed 0–10 rubric — none is ever from the maker’s own lab, and an answer that breaks a task’s explicit rules is capped in code, not by the judge’s goodwill. The score is the average, out of 100.
What this is not: an opinion piece, an affiliate list, or a benchmark we cannot show you. Every score links its raw outputs and judge verdicts. The full protocol · receipts: GPT-5.6 Sol, GPT-5.6 Terra, GPT-5.5, Claude Opus 4.6, GPT-5.6 Luna, Kimi K3, GLM 5.2, Gemini 3.1 Pro Preview, GPT-5.3-Codex, Grok 4.5, Qwen3.7 Max, Claude Opus 4.8, DeepSeek V4 Pro, Claude Fable 5, DeepSeek V4 Flash, Gemini 3.5 Flash, Claude Sonnet 5, Gemini 3.1 Flash Lite, Mistral Medium 3.5
Questions people ask
What is the best AI for revision and study notes in 2026?
GPT-5.6 Sol leads our tested ranking with 94/100 on our 12-task study suite (12 tasks), in a dead heat with GPT-5.6 Terra (94). Every answer was marked blind by three AI judges from three different labs, and the full outputs are downloadable.
How is this ranking made?
Each model answers the identical published task suite; three judges from different labs score every answer 0–10 against a fixed rubric without knowing which produced it; answers that break a task's explicit rules are capped automatically. The score is the average, out of 100. No vendor pays for placement.
How often does this page update?
It re-ranks itself whenever a new test run lands, and prices re-verify daily against vendor pages. The current ranking was last computed on 14 Aug 2026.
More rankings: Best AI for writing · Best AI chatbot for everyday use · Best AI for coding · Best free AI model · Best AI for spreadsheets and Excel · Best AI for writing essays · Best AI for summarising documents · Best AI for extracting data from text · Best AI for writing emails · Best AI for everyday maths and percentages · Best AI for customer service replies · Best-value AI model API · every model we track · every tool