Best AI for presentations
Scored 98/100 on our 12-task presentation suite — level on quality with GPT-5.5 (98) — the top spot goes on the tie-break: cleanest rule-compliance, then lowest measured cost; GPT-5.3-Codex sits a single point behind.
GPT-5.6 Luna is the engine inside ChatGPT. Go to chatgpt.com ↗ — the free tier is fine to start. Paid plans start at £7/month (go plan, vendor’s own price). The free tier is ChatGPT’s, not a promise about this exact model — we haven’t verified which plan carries it. Not fussed about the last point or two? Any of the top 3 here will serve you well.
A presentation is a talk with pictures, not a document read aloud — so the suite tests both halves: cutting a nine-bullet slide to the four that help anyone decide, speaker notes that sound like a person, an opening that earns 30 seconds of a sixth-former's attention, and preparing for the Q&A nobody enjoys. One task asks for an investor slide claiming 50,000 users when the truth is 3,100 — the right answer is to refuse.
updated 16 Aug 2026 · tested by Robert Prime · re-ranks automatically when a new run lands
| # | Model | Our score |
|---|---|---|
| 1 | GPT-5.6 Lunalatest | 98/100 |
| 2 | GPT-5.5 | 98/100 |
| 3 | GPT-5.3-Codex | 97/100 |
| 4 | GPT-5.6 Terra | 96/100 |
| 5 | GPT-5.6 Sol | 95/100 |
| 6 | Claude Opus 4.8 | 92/100 |
| 7 | Claude Fable 5 | 92/100 |
| 8 | DeepSeek V4 Pro | 91/100 |
| 9 | Grok 4.5 | 91/100 |
| 10 | Qwen3.7 Max | 89/100 |
| 11 | Claude Sonnet 5 | 88/100 |
| 12 | GLM 5.2 | 87/100 |
| 13 | Gemini 3.5 Flash | 87/100 |
| 14 | Mistral Medium 3.5 | 86/100 |
| 15 | Kimi K3 | 86/100 |
| 16 | Claude Opus 4.6 | 85/100 |
| 17 | DeepSeek V4 Flash | 84/100 |
| 18 | Gemini 3.1 Pro Preview | 83/100 |
| 19 | Gemini 3.1 Flash Lite | 82/100 |
“API cost” is what software developers pay to build on a model — ignore it if you just use the website. Each model answers each task once. Models level on score are ranked by a fixed tie-break — fewest machine-checked rule breaches, then lowest measured cost per run — so the order is deterministic and checkable, never arbitrary. Judge panels never include the contestant’s own lab, so panels differ slightly per model — small cross-model gaps can reflect panel severity, not quality.
Made by OpenAI — their newest model. You use it inside ChatGPT — nothing to install.
Strongest showing: “Same facts, different room” — scored 10/10 by the panel. Weakest: “One slide, not a memo” at 9/10.
“The response perfectly adapts the facts for a Year 6 audience, includes a clear and concrete call to action, and strictly adheres to the word count limit.”— google/gemini-3.1-pro-preview, judging blind · full receipts ↓
12 tasks · 16 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $0.1 in / $0.6 out per 1M tokens · full model page →
Same maker, same score — effectively the same choice as #1; the deep-dive above covers it. Full model page →
Made by OpenAI. You use it inside ChatGPT — nothing to install.
Strongest showing: “Hostile Q&A prep” — scored 10/10 by the panel. Weakest: “One slide, not a memo” at 9/10.
“The response perfectly follows all instructions, providing exactly 5 hard, realistic questions with honest, one-sentence answers on the same line separated by a dash. No constraints were violated.”— google/gemini-3.1-pro-preview, judging blind · full receipts ↓
12 tasks · 16 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $1.75 in / $14 out per 1M tokens · full model page →
See it for yourself
One task from the suite, and what the leader, the middle of the table and the bottom actually wrote. This is the task they disagreed on most — scores ranged from 5 to 10 out of 10.
Write the spoken opening (max 80 words) of a careers-evening talk to 40 sixth-formers about plumbing apprenticeships. It must earn attention in the first sentence with something concrete from these true facts: day-one pay £15k rising to £32k qualified; a Christmas Day callout; zero student debt. It must not begin with 'Hello' or 'Good evening' and must not introduce the speaker — banned phrase: 'my name is'.
On day one of a plumbing apprenticeship, you can earn £15,000—and, once qualified, that can rise to £32,000. You could be called out on Christmas Day, solving a problem when a family really needs you. And unlike university, you can build a skilled career without student debt. Tonight, let’s look at…
Judge: The response perfectly follows all instructions. It meets the word count, avoids the banned greetings and phrases, and immediately grabs attention using the provided facts. The tone is engaging and highly appropriate for the target audience
Imagine earning £15k on day one, graduating with zero student debt, and eventually cashing in on a Christmas Day callout on your way to a £32k qualified salary. That’s the reality of a plumbing apprenticeship. While your peers stress over university tuition, you get paid to learn and walk into a car…
Judge: Meets length, opening-hook, banned-phrase, and speaker-introduction constraints. Uses the supplied facts effectively and is engaging. Minor issue: “completely secure your financial future” is an overstatement not supported by the facts.
Imagine waking up on Christmas morning to a frantic call about a burst pipe, fixing it by noon, and pocketing a triple-time bonus before the turkey is even carved. While your friends are racking up fifty grand in student debt for a degree they might never use, you could be earning fifteen thousand p…
Judge: Capped in code — the answer broke a checkable rule (81 words against a maximum of 80). Strong, engaging opening using the specified pay, Christmas callout and debt themes. However, it is 81 words, exceeding the 80-word maximum. “Triple-time bonus” and £50k debt are unsupported additions.
Trimmed to fit. Every full answer, every judge’s score: GPT-5.6 Luna, Qwen3.7 Max, Gemini 3.1 Flash Lite
How this ranking is made
Every model answers the identical published task suite. Three AI judges from three different labs mark each answer blind against a fixed 0–10 rubric — none is ever from the maker’s own lab, and an answer that breaks a task’s explicit rules is capped in code, not by the judge’s goodwill. The score is the average, out of 100.
What this is not: an opinion piece, an affiliate list, or a benchmark we cannot show you. Every score links its raw outputs and judge verdicts. The full protocol · receipts: GPT-5.6 Luna, GPT-5.5, GPT-5.3-Codex, GPT-5.6 Terra, GPT-5.6 Sol, Claude Opus 4.8, Claude Fable 5, DeepSeek V4 Pro, Grok 4.5, Qwen3.7 Max, Claude Sonnet 5, GLM 5.2, Gemini 3.5 Flash, Mistral Medium 3.5, Kimi K3, Claude Opus 4.6, DeepSeek V4 Flash, Gemini 3.1 Pro Preview, Gemini 3.1 Flash Lite
Questions people ask
What is the best AI for presentations in 2026?
GPT-5.6 Luna leads our tested ranking with 98/100 on our 12-task presentation suite (12 tasks), in a dead heat with GPT-5.5 (98). Every answer was marked blind by three AI judges from three different labs, and the full outputs are downloadable.
How is this ranking made?
Each model answers the identical published task suite; three judges from different labs score every answer 0–10 against a fixed rubric without knowing which produced it; answers that break a task's explicit rules are capped automatically. The score is the average, out of 100. No vendor pays for placement.
How often does this page update?
It re-ranks itself whenever a new test run lands, and prices re-verify daily against vendor pages. The current ranking was last computed on 16 Aug 2026.
Head-to-heads in presentations
Show all 20 tested pairs ▾
More rankings ▾
Best AI for writing · Best AI chatbot for everyday use · Best AI for coding · Best free AI model · Best AI for spreadsheets and Excel · Best AI essay writer · Best AI for summarising documents · Best AI for extracting data from text · Best AI for writing emails · Best AI for everyday maths and percentages · Best AI for customer service replies · Best AI for revision and study notes · Best AI for vibe coding · Best AI for making flashcards · Best AI for social media posts · Best AI for job applications and cover letters · Best AI for writing your CV · Best AI for research skills · Best AI for creative writing · Best AI for translation · Best AI for travel planning · Best AI for emotional support · Best AI for everyday legal questions · Best AI for health questions · Best AI for writing a book · Best AI for humanising AI text · Best-value AI model API · every model we track · every tool