Best for / Maths
Best AI for everyday maths and percentages
Not maths olympiad — the numbers real life throws at you: VAT backwards, the discount that stacks, the reverse percentage every shop hopes you get wrong. Every answer is checkable arithmetic, and the traps are the classic human mistakes, so a right answer means it will not walk you into them. One task demands an 'exact' 10-year investment figure; the honest answer is a range with assumptions.
updated 14 Aug 2026 · tested by Robert Prime · re-ranks automatically when a new run lands
Scored 100/100 on our 12-task everyday-maths suite — in a dead heat with GPT-5.6 Sol (99), so either is a fine choice.
| # | Model | Our score | Context | Free? | API $/1M in·out |
|---|---|---|---|---|---|
| 1 | Claude Opus 4.8 | 100/100 | 1000k | — | $5 · $25 |
| 2 | GPT-5.6 Sollatest | 99/100 | 1050k | — | $5 · $30 |
| 3 | GPT-5.6 Terralatest | 99/100 | 1050k | — | $1 · $6 |
| 4 | Claude Opus 4.6 | 98/100 | 1000k | — | $5 · $25 |
| 5 | GPT-5.3-Codex | 98/100 | 400k | — | $1.75 · $14 |
| 6 | GPT-5.5 | 98/100 | 1050k | — | $5 · $30 |
| 7 | GPT-5.6 Lunalatest | 98/100 | 1050k | — | $0.1 · $0.6 |
| 8 | Grok 4.5latest | 98/100 | 500k | grok.com | $2 · $6 |
| 9 | Kimi K3latest | 98/100 | 1049k | — | $3 · $15 |
| 10 | GLM 5.2latest | 97/100 | 1049k | chat.z.ai | $0.63 · $1.98 |
| 11 | Claude Fable 5 | 96/100 | 1000k | — | $10 · $50 |
| 12 | DeepSeek V4 Prolatest | 96/100 | 1049k | — | $1.168 · $2.336 |
| 13 | Gemini 3.1 Flash Lite | 96/100 | 1049k | — | $0.25 · $1.5 |
| 14 | Gemini 3.1 Pro Preview | 96/100 | 1049k | — | $2 · $12 |
| 15 | Gemini 3.5 Flash | 96/100 | 1049k | Google AI Studio | $1.5 · $9 |
| 16 | Qwen3.7 Maxlatest | 96/100 | 1000k | — | $1.475 · $4.425 |
| 17 | DeepSeek V4 Flashlatest | 93/100 | 1049k | — | $0.14 · $0.28 |
| 18 | Claude Sonnet 5latest | 90/100 | 1000k | — | $2 · $10 |
| 19 | Mistral Medium 3.5latest | 74/100 | 262k | — | $1.5 · $7.5 |
anthropic's model, holding about 1000k tokens of context — reached through the API.
Strongest showing: “VAT both ways” — scored 10/10 by the panel.
“The response accurately calculates the ex-VAT price and VAT amount, showing each calculation on a single line as requested. It correctly avoids the common mistake and provides a helpful, concise explanation.”— google/gemini-3.1-pro-preview, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5 · API $5 in / $25 out per 1M tokens · full model page →
openai's newest model, holding about 1050k tokens of context — reached through the API.
Strongest showing: “VAT both ways” — scored 10/10 by the panel. Weakest: “Mortgage overpayment intuition” at 9/10.
“Correct calculation, avoids common error, concise one-line format as requested, clear and useful.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $5 in / $30 out per 1M tokens · full model page →
openai's newest model, holding about 1050k tokens of context — reached through the API.
Strongest showing: “VAT both ways” — scored 10/10 by the panel. Weakest: “Rate to reality” at 9/10.
“Correct calculation, avoids the common error, concise one-line format as requested, clear for non-technical reader.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $1 in / $6 out per 1M tokens · full model page →
anthropic's model, holding about 1000k tokens of context — reached through the API.
Strongest showing: “VAT both ways” — scored 10/10 by the panel. Weakest: “Refuse false precision” at 9/10.
“The response perfectly calculates the ex-VAT price and VAT amount, showing the correct calculations in one line each as requested, and avoids the common mathematical error.”— google/gemini-3.1-pro-preview, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5 · API $5 in / $25 out per 1M tokens · full model page →
openai's model, holding about 400k tokens of context — reached through the API.
Strongest showing: “VAT both ways” — scored 10/10 by the panel. Weakest: “Estimate honestly” at 9/10. The judges flagged 1 instruction breach across the run — those scores were capped automatically.
“Correct calculation, avoids the common error, concise one-line format as requested.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $1.75 in / $14 out per 1M tokens · full model page →
openai's model, holding about 1050k tokens of context — reached through the API.
Strongest showing: “VAT both ways” — scored 10/10 by the panel. Weakest: “Refuse false precision” at 9/10. The judges flagged 1 instruction breach across the run — those scores were capped automatically.
“Correct calculations, avoids common error, one-line each as instructed, clear and concise.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $5 in / $30 out per 1M tokens · full model page →
openai's newest model, holding about 1050k tokens of context — reached through the API.
Strongest showing: “VAT both ways” — scored 10/10 by the panel. Weakest: “Estimate honestly” at 9/10.
“Correct calculations, avoids the common error, concise one-line format as requested.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $0.1 in / $0.6 out per 1M tokens · full model page →
x-ai's newest model, holding about 500k tokens of context — reached through the API or free via grok.com.
Strongest showing: “VAT both ways” — scored 10/10 by the panel. Weakest: “Mortgage overpayment intuition” at 8/10. The judges flagged 1 instruction breach across the run — those scores were capped automatically.
“Correct calculation, avoids the common error, follows format with one-line calculations, clear and concise.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $2 in / $6 out per 1M tokens · free via grok.com ↗ (checked 11 Aug 2026) · full model page →
moonshotai's newest model, holding about 1049k tokens of context — reached through the API.
Strongest showing: “VAT both ways” — scored 10/10 by the panel. Weakest: “Refuse false precision” at 9/10. The judges flagged 2 instruction breaches across the run — those scores were capped automatically.
“Correct calculation, correct method avoiding the trap, clear one-line format with brief explanation. Accurate and concise.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $3 in / $15 out per 1M tokens · full model page →
z-ai's newest model, holding about 1049k tokens of context — reached through the API or free via chat.z.ai.
Strongest showing: “VAT both ways” — scored 10/10 by the panel. Weakest: “Refuse false precision” at 8/10.
“Correct calculation, avoids the common error, follows one-line format instructions, concise and clear.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $0.63 in / $1.98 out per 1M tokens · free via chat.z.ai ↗ (checked 11 Aug 2026) · full model page →
anthropic's model, holding about 1000k tokens of context — reached through the API.
Strongest showing: “VAT both ways” — scored 10/10 by the panel. Weakest: “Percentage change vs points” at 5/10. The judges flagged 4 instruction breaches across the run — those scores were capped automatically.
“The response perfectly calculates the ex-VAT price and VAT amount, showing each calculation on a single line as requested. It correctly avoids the common pitfall and provides a clear, concise explanation.”— google/gemini-3.1-pro-preview, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5 · API $10 in / $50 out per 1M tokens · full model page →
deepseek's newest model, holding about 1049k tokens of context — reached through the API.
Strongest showing: “VAT both ways” — scored 10/10 by the panel. Weakest: “Split a bill with a twist” at 5/10. The judges flagged 3 instruction breaches across the run — those scores were capped automatically.
“Correct calculation, avoids common error, concise one-liners, meets task requirements accurately.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $1.168 in / $2.336 out per 1M tokens · full model page →
google's model, holding about 1049k tokens of context — reached through the API.
Strongest showing: “Percentage change vs points” — scored 10/10 by the panel. Weakest: “Estimate honestly” at 8/10. The judges flagged 1 instruction breach across the run — those scores were capped automatically.
“Correct figures (1pp, 25%), meets task in two sentences, clear and concise, minor lack of explicit labeling but accurate and useful.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gpt-5.6-terra, grok-4.5 · API $0.25 in / $1.5 out per 1M tokens · full model page →
google's model, holding about 1049k tokens of context — reached through the API.
Strongest showing: “VAT both ways” — scored 10/10 by the panel. Weakest: “Refuse false precision” at 8/10. The judges flagged 1 instruction breach across the run — those scores were capped automatically.
“Correct calculations, avoids the common error, follows one-line-each format, clear and concise.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gpt-5.6-terra, grok-4.5 · API $2 in / $12 out per 1M tokens · full model page →
google's model, holding about 1049k tokens of context — reached through the API or free via Google AI Studio.
Strongest showing: “VAT both ways” — scored 10/10 by the panel. Weakest: “Estimate honestly” at 8/10.
“Correct calculations, correct approach avoiding the common error, concise one-line format as requested.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gpt-5.6-terra, grok-4.5 · API $1.5 in / $9 out per 1M tokens · free via Google AI Studio ↗ (checked 11 Aug 2026) · full model page →
qwen's newest model, holding about 1000k tokens of context — reached through the API.
Strongest showing: “VAT both ways” — scored 10/10 by the panel. Weakest: “Refuse false precision” at 8/10.
“Correct calculation, avoids the common error, concise one-line format as requested.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $1.475 in / $4.425 out per 1M tokens · full model page →
deepseek's newest model, holding about 1049k tokens of context — reached through the API.
Strongest showing: “VAT both ways” — scored 10/10 by the panel. Weakest: “Percentage change vs points” at 5/10. The judges flagged 2 instruction breaches across the run — those scores were capped automatically.
“Correct calculation, avoids common error, concise one-line format as requested.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $0.14 in / $0.28 out per 1M tokens · full model page →
anthropic's newest model, holding about 1000k tokens of context — reached through the API.
Strongest showing: “VAT both ways” — scored 10/10 by the panel. Weakest: “Percentage change vs points” at 5/10. The judges flagged 3 instruction breaches across the run — those scores were capped automatically.
“The response accurately calculates the ex-VAT price and VAT amount, shows the calculations in one line each as requested, and provides a clear, helpful explanation of why the common mistake is wrong.”— google/gemini-3.1-pro-preview, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5 · API $2 in / $10 out per 1M tokens · full model page →
mistralai's newest model, holding about 262k tokens of context — reached through the API.
Strongest showing: “VAT both ways” — scored 10/10 by the panel. Weakest: “Rate to reality” at 3/10. The judges flagged 6 instruction breaches across the run — those scores were capped automatically.
“Correct calculation, avoids the common error, concise one-line format as requested.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $1.5 in / $7.5 out per 1M tokens · full model page →
See it for yourself
One task from the suite, and what the leader, the middle of the table and the bottom actually wrote. This is the task they disagreed on most — scores ranged from 3 to 10 out of 10.
A £80 item has '25% off' and then a further '10% off at checkout'. What is the final price, and why is it not 35% off? Two sentences maximum plus the number.
The final price is £54. It's not 35% off because the 10% discount applies to the already-reduced price of £60 (not the original £80), so you get 90% of 75% of the price, which equals 67.5% of the original—a total discount of 32.5%.
Judge: The response perfectly calculates the final price, accurately explains the math behind successive discounts, and strictly adheres to the two-sentence maximum constraint.
£54 The second 10% discount is applied to the already-reduced price of £60, not the original £80. Because the discounts compound sequentially on the new subtotal, the total reduction is 32.5% rather than a flat 35%.
Judge: Correct math (80*0.75*0.9=54), correct explanation, within two-sentence plus number constraint though number given first, minor formatting deviation but content accurate and clear.
The final price is **£52** because the discounts are applied sequentially (25% off first, then 10% off the reduced price), not added together. Sequential discounts multiply, not add. £52
Judge: Incorrect final price: correct answer is £54, not £52. Also redundantly repeats the number, slightly violating the format instruction.
Trimmed to fit. Every full answer, every judge’s score: Claude Opus 4.8, GLM 5.2, Mistral Medium 3.5
How this ranking is made
Every model answers the identical published task suite. Three AI judges from three different labs mark each answer blind against a fixed 0–10 rubric — none is ever from the maker’s own lab, and an answer that breaks a task’s explicit rules is capped in code, not by the judge’s goodwill. The score is the average, out of 100.
What this is not: an opinion piece, an affiliate list, or a benchmark we cannot show you. Every score links its raw outputs and judge verdicts. The full protocol · receipts: Claude Opus 4.8, GPT-5.6 Sol, GPT-5.6 Terra, Claude Opus 4.6, GPT-5.3-Codex, GPT-5.5, GPT-5.6 Luna, Grok 4.5, Kimi K3, GLM 5.2, Claude Fable 5, DeepSeek V4 Pro, Gemini 3.1 Flash Lite, Gemini 3.1 Pro Preview, Gemini 3.5 Flash, Qwen3.7 Max, DeepSeek V4 Flash, Claude Sonnet 5, Mistral Medium 3.5
Questions people ask
What is the best AI for everyday maths and percentages in 2026?
Claude Opus 4.8 leads our tested ranking with 100/100 on our 12-task everyday-maths suite (12 tasks), in a dead heat with GPT-5.6 Sol (99). Every answer was marked blind by three AI judges from three different labs, and the full outputs are downloadable.
How is this ranking made?
Each model answers the identical published task suite; three judges from different labs score every answer 0–10 against a fixed rubric without knowing which produced it; answers that break a task's explicit rules are capped automatically. The score is the average, out of 100. No vendor pays for placement.
How often does this page update?
It re-ranks itself whenever a new test run lands, and prices re-verify daily against vendor pages. The current ranking was last computed on 14 Aug 2026.
More rankings: Best AI for writing · Best AI chatbot for everyday use · Best AI for coding · Best free AI model · Best AI for spreadsheets and Excel · Best AI for writing essays · Best AI for summarising documents · Best AI for extracting data from text · Best AI for writing emails · Best AI for customer service replies · Best AI for revision and study notes · Best-value AI model API · every model we track · every tool