Best AI for meeting notes
Scored 94/100 on our 12-task meetings suite — a single point ahead of GPT-5.6 Sol (93) — effectively level. Free to use via grok.com.
Grok 4.5 is the engine inside Grok. Go to grok.com ↗ — free to use. Not fussed about the last point or two? Any of the top 2 here will serve you well.
Every task here hands the model a real-shaped transcript — with the crosstalk, the corrections, the person who changes their mind and the action item nobody picks up — and asks for the notes. The thing being tested is restraint: the failure that makes AI notetakers untrustworthy is not missing a point, it is inventing a decision. One task contains a discussion where NO decision was reached, and a model that records one has failed it however well written the rest is. Two tasks have exactly one correct answer and are marked by machine, not by opinion: who finally owned each action, and what the relative dates resolve to.
updated 28 Aug 2026 · tested by Robert Prime · re-ranks automatically when a new run lands
- All 19 sat the identical 12-task suite — same tasks, same order, one attempt each.
- Every answer marked blind: judges are not told which entrant wrote it.
- Three judges per answer, each from a competing lab. A panel never includes the entrant's own lab.
- Scored 0–10 against a fixed rubric. Answers breaking a task's explicit rules are capped by machine, not by opinion.
- Ties broken by fewest rule breaches, then lowest measured cost — published, not editorial.
- Nobody pays for placement. Last computed 28 Aug 2026.
Every score below links to the raw file behind it: every task, the entrant’s real answer, and all three judges’ marks.
| # | Model | Our score |
|---|---|---|
| 1 | Grok 4.5latest | 94/100 |
| 2 | GPT-5.6 Sollatest | 93/100 |
| 3 | GPT-5.5 | 92/100 |
| 4 | GPT-5.3-Codex | 91/100 |
| 5 | Qwen3.7 Max | 89/100 |
| 6 | Claude Opus 4.6 | 88/100 |
| 7 | GPT-5.6 Terra | 88/100 |
| 8 | GLM 5.2 | 88/100 |
| 9 | Claude Opus 4.8 | 88/100 |
| 10 | Claude Sonnet 5 | 87/100 |
| 11 | Kimi K3 | 87/100 |
| 12 | DeepSeek V4 Pro | 86/100 |
| 13 | Gemini 3.5 Flash | 86/100 |
| 14 | GPT-5.6 Luna | 86/100 |
| 15 | Gemini 3.1 Pro Preview | 86/100 |
| 16 | Mistral Medium 3.5 | 84/100 |
| 17 | DeepSeek V4 Flash | 81/100 |
| 18 | Claude Fable 5 | 79/100 |
| 19 | Gemini 3.1 Flash Lite | 77/100 |
“API cost” is what software developers pay to build on a model — ignore it if you just use the website. Each model answers each task once. Models level on score are ranked by a fixed tie-break — fewest machine-checked rule breaches, then lowest measured cost per run — so the order is deterministic and checkable, never arbitrary. Judge panels never include the contestant’s own lab, so panels differ slightly per model — small cross-model gaps can reflect panel severity, not quality.
Made by xAI — their newest model. You use it inside Grok — nothing to install; free via grok.com.
Strongest showing: “Actions, owners, and the one with neither” — scored 10/10 by the panel. Weakest: “Who has not spoken” at 8/10. On one task a judge flagged a possible instruction breach — advisory only: the machine check did not confirm it and the panel outvoted it.
“All four values correct, final ownership properly resolved, exact JSON format with required keys only.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 28 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $2 in / $6 out per 1M tokens · free via grok.com ↗ (checked 11 Aug 2026) · full model page →
Made by OpenAI — their newest model. You use it inside ChatGPT — nothing to install.
Strongest showing: “Notes for the person who missed it” — scored 9/10 by the panel. Weakest: “Actions, owners, and the one with neither” at 5/10. Our automated rule-check caught 1 instruction breach across the run (a word limit or banned phrase) — capped in code, separately from the judges’ marks.
“Leads with marketing lead's action on rebrand, retains relevant headcount impact, drops irrelevant items. Exactly 3 bullets, under 120 words, clear and useful.”— x-ai/grok-4.5, judging blind · full receipts ↓
12 tasks · 28 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $2 in / $10 out per 1M tokens · full model page →
Made by OpenAI. You use it inside ChatGPT — nothing to install.
Strongest showing: “The number said twice” — scored 10/10 by the panel. Weakest: “Actions, owners, and the one with neither” at 5/10. Our automated rule-check caught 1 instruction breach across the run (a word limit or banned phrase) — capped in code, separately from the judges’ marks.
“The response perfectly follows all instructions, providing a clear, accurate, three-line summary under the word limit without including the retracted figure.”— google/gemini-3.1-pro-preview, judging blind · full receipts ↓
12 tasks · 28 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $5 in / $30 out per 1M tokens · full model page →
See it for yourself
One task from the suite, and what the leader, the middle of the table and the bottom actually wrote. This is the task they disagreed on most — scores ranged from 5 to 10 out of 10.
Transcript extract. MAYA: 'I'll get the pricing page updated before the campaign goes out.' TOM: 'And someone needs to tell support about the new tiers.' MAYA: 'Yeah.' TOM: 'I can do the deck for Thursday.' MAYA: 'Perfect. Oh and the legal review — Ash said he'd look but that was three weeks ago.' TOM: 'I'll chase Ash.' MAYA: 'Actually leave it, I'll chase him, I owe him an email anyway.' Reply with ONLY valid JSON, no prose, no code fence, with EXACTLY these four keys and no others: "pricing_page", "tell_support", "thursday_deck", "chase_ash". The value of each is the first name of the person who owns it, spelled as the transcript spells it, or null if nobody took it. Where ownership changed during the conversation, give the FINAL owner.
{"pricing_page": "Maya", "tell_support": null, "thursday_deck": "Tom", "chase_ash": "Maya"}Judge: All four values correct, final ownership properly resolved, exact JSON format with required keys only.
{"pricing_page": "Maya", "tell_support": null, "thursday_deck": "Tom", "chase_ash": "Maya"}Judge: The response perfectly follows all instructions, extracting the correct final owners for each task and formatting the output exactly as requested without code fences or prose.
{
"pricing_page": "MAYA",
"tell_support": null,
"thursday_deck": "TOM",
"chase_ash": "MAYA"
}Judge: Capped in code — the answer broke a checkable rule (wrong answer at $.pricing_page: expected "Maya", got "MAYA"). Names must be spelled as first names (Maya, Tom), not all caps as in response; transcript spells names capitalized only as speaker labels, first-name spelling should be 'Maya'/'Tom' not 'MAYA'/'TOM'.
Trimmed to fit. Every full answer, every judge’s score: Grok 4.5, Claude Sonnet 5, Gemini 3.1 Flash Lite
How this ranking is made
Every model answers the identical published task suite. Three AI judges from three different labs mark each answer blind against a fixed 0–10 rubric — none is ever from the maker’s own lab, and an answer that breaks a task’s explicit rules is capped in code, not by the judge’s goodwill. The score is the average, out of 100.
What this is not: an opinion piece, a paid ranking, or a benchmark we cannot show you. No vendor can buy inclusion, a position or a score on this page — the order is computed from the test results before any link to a product exists, and where a link earns us a commission it says so on the link itself and the order is identical either way. Every score links its raw outputs and judge verdicts. The full protocol · How we make money · receipts: Grok 4.5, GPT-5.6 Sol, GPT-5.5, GPT-5.3-Codex, Qwen3.7 Max, Claude Opus 4.6, GPT-5.6 Terra, GLM 5.2, Claude Opus 4.8, Claude Sonnet 5, Kimi K3, DeepSeek V4 Pro, Gemini 3.5 Flash, GPT-5.6 Luna, Gemini 3.1 Pro Preview, Mistral Medium 3.5, DeepSeek V4 Flash, Claude Fable 5, Gemini 3.1 Flash Lite
Questions people ask
What is the best AI for meeting notes in 2026?
Grok 4.5 leads our tested ranking with 94/100 on our 12-task meetings suite, in a dead heat with GPT-5.6 Sol (93). Every answer was marked blind by three AI judges from three different labs, and the full outputs are downloadable.
How is this ranking made?
Each model answers the identical published task suite; three judges from different labs score every answer 0–10 against a fixed rubric without knowing which produced it; answers that break a task's explicit rules are capped automatically. The score is the average, out of 100. No vendor pays for placement.
What happens when two models score the same?
They are separated by a fixed, published tie-break rather than by editorial choice: first the fewest machine-checked rule breaches, then the lowest measured cost per run. The order is therefore reproducible from the same data, and where the tie-break decides the top place the page says so.
How often does this page update?
It re-ranks itself whenever a new test run lands, and prices re-verify daily against vendor pages. The current ranking was last computed on 28 Aug 2026.
Is Grok 4.5 free?
Yes — Grok 4.5 can be used at no cost via grok.com, confirmed on 11 Aug 2026.
Head-to-heads in meetings & notes
Show all 20 tested pairs ▾
More rankings ▾
Best AI for writing · Best AI chatbot for everyday use · Best AI for coding · Best free AI model · Best AI for spreadsheets and Excel · Best AI essay writer · Best AI for summarising documents · Best AI for extracting data from text · Best AI for writing emails · Best AI for everyday maths and percentages · Best AI for customer service replies · Best AI for revision and study notes · Best AI for vibe coding · Best AI for making flashcards · Best AI for social media posts · Best AI for job applications and cover letters · Best AI for presentations · Best AI for writing your CV · Best AI for research skills · Best AI for creative writing · Best AI for translation · Best AI for travel planning · Best AI for emotional support · Best AI for everyday legal questions · Best AI for health questions · Best AI for writing a book · Best AI for humanising AI text · Best AI research agent · Best AI for code review · Best AI for workflow automation · Best AI for bookkeeping · Best AI for HR and employment questions · Best AI for property and lettings · Best-value AI model API · every model we track · every tool