Best AI for everyday maths and percentages

The verdict

Claude Opus 4.8

Scored 100/100 on our 12-task everyday-maths suitea single point ahead of GPT-5.6 Terra (99) — effectively level.

What to actually do

Claude Opus 4.8 is the engine inside Claude. Go to claude.aithe free tier is fine to start. Paid plans start at $20/month (pro plan, vendor’s own price). The free tier is Claude’s, not a promise about this exact model — we haven’t verified which plan carries it. Not fussed about the last point or two? Any of the top 3 here will serve you well.

Not maths olympiad — the numbers real life throws at you: VAT backwards, the discount that stacks, the reverse percentage every shop hopes you get wrong. Every answer is checkable arithmetic, and the traps are the classic human mistakes, so a right answer means it will not walk you into them. One task demands an 'exact' 10-year investment figure; the honest answer is a range with assumptions.

updated 14 Aug 2026 · tested by Robert Prime · re-ranks automatically when a new run lands

How this was measured
  • All 19 sat the identical 12-task suite — same tasks, same order, one attempt each.
  • Every answer marked blind: judges are not told which entrant wrote it.
  • Three judges per answer, each from a competing lab. A panel never includes the entrant's own lab.
  • Scored 0–10 against a fixed rubric. Answers breaking a task's explicit rules are capped by machine, not by opinion.
  • Ties broken by fewest rule breaches, then lowest measured cost — published, not editorial.
  • Nobody pays for placement. Last computed 14 Aug 2026.

Every score below links to the raw file behind it: every task, the entrant’s real answer, and all three judges’ marks.

#ModelOur score
1Claude Opus 4.8100/100
2GPT-5.6 Terralatest99/100
3GPT-5.6 Sollatest99/100
4GPT-5.6 Luna98/100
5Grok 4.598/100
6Claude Opus 4.698/100
7GPT-5.3-Codex98/100
8GPT-5.598/100
9Kimi K398/100
10GLM 5.297/100
11Gemini 3.1 Flash Lite96/100
12DeepSeek V4 Pro96/100
13Gemini 3.5 Flash96/100
14Qwen3.7 Max96/100
15Gemini 3.1 Pro Preview96/100
16Claude Fable 596/100
17DeepSeek V4 Flash93/100
18Claude Sonnet 590/100
19Mistral Medium 3.574/100

“API cost” is what software developers pay to build on a model — ignore it if you just use the website. Each model answers each task once. Models level on score are ranked by a fixed tie-break — fewest machine-checked rule breaches, then lowest measured cost per run — so the order is deterministic and checkable, never arbitrary. Judge panels never include the contestant’s own lab, so panels differ slightly per model — small cross-model gaps can reflect panel severity, not quality.

1.

Claude Opus 4.8

100/100our pick — won on the tie-break

Made by Anthropic. You use it inside Claude — nothing to install.

Strongest showing: Percentage change vs points” — scored 10/10 by the panel.

The response accurately calculates both values, follows the two-sentence constraint perfectly, and includes both required numbers. It is clear, concise, and highly useful.google/gemini-3.1-pro-preview, judging blind · full receipts ↓

12 tasks · 14 Aug 2026 · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5 · API $5 in / $25 out per 1M tokens · full model page →

Made by OpenAI — their newest model. You use it inside ChatGPT — nothing to install.

Strongest showing: Percentage change vs points” — scored 10/10 by the panel. Weakest: “Rate to reality” at 9/10.

The response perfectly follows all instructions, providing accurate calculations in exactly one sentence each, and includes both numbers as requested. It is clear and concise.google/gemini-3.1-pro-preview, judging blind · full receipts ↓

12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $2 in / $12 out per 1M tokens · full model page →

3.

GPT-5.6 Sol

99/100

Made by OpenAI — their newest model. You use it inside ChatGPT — nothing to install.

Strongest showing: Percentage change vs points” — scored 10/10 by the panel. Weakest: “Mortgage overpayment intuition” at 9/10.

The response accurately calculates both values, includes both numbers, and perfectly follows the constraint of using exactly one sentence for each explanation.google/gemini-3.1-pro-preview, judging blind · full receipts ↓

12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $2 in / $10 out per 1M tokens · full model page →

See it for yourself

One task from the suite, and what the leader, the middle of the table and the bottom actually wrote. This is the task they disagreed on most — scores ranged from 3 to 10 out of 10.

The task — The discount stack

A £80 item has '25% off' and then a further '10% off at checkout'. What is the final price, and why is it not 35% off? Two sentences maximum plus the number.

The final price is £54. It's not 35% off because the 10% discount applies to the already-reduced price of £60 (not the original £80), so you get 90% of 75% of the price, which equals 67.5% of the original—a total discount of 32.5%.

Judge: The response perfectly calculates the final price, accurately explains the math behind successive discounts, and strictly adheres to the two-sentence maximum constraint.

GLM 5.210/10
£54

The second 10% discount is applied to the already-reduced price of £60, not the original £80. Because the discounts compound sequentially on the new subtotal, the total reduction is 32.5% rather than a flat 35%.

Judge: The response perfectly calculates the final price, accurately explains the sequential discounting in exactly two sentences, and includes the final number as requested.

The final price is **£52** because the discounts are applied sequentially (25% off first, then 10% off the reduced price), not added together. Sequential discounts multiply, not add. £52

Judge: The response contains a major calculation error; the final price should be £54, not £52. It followed the sentence constraint but failed on basic accuracy.

Trimmed to fit. Every full answer, every judge’s score: Claude Opus 4.8, GLM 5.2, Mistral Medium 3.5

How this ranking is made

Every model answers the identical published task suite. Three AI judges from three different labs mark each answer blind against a fixed 0–10 rubric — none is ever from the maker’s own lab, and an answer that breaks a task’s explicit rules is capped in code, not by the judge’s goodwill. The score is the average, out of 100.

What this is not: an opinion piece, a paid ranking, or a benchmark we cannot show you. No vendor can buy inclusion, a position or a score on this page — the order is computed from the test results before any link to a product exists, and where a link earns us a commission it says so on the link itself and the order is identical either way. Every score links its raw outputs and judge verdicts. The full protocol · How we make money · receipts: Claude Opus 4.8, GPT-5.6 Terra, GPT-5.6 Sol, GPT-5.6 Luna, Grok 4.5, Claude Opus 4.6, GPT-5.3-Codex, GPT-5.5, Kimi K3, GLM 5.2, Gemini 3.1 Flash Lite, DeepSeek V4 Pro, Gemini 3.5 Flash, Qwen3.7 Max, Gemini 3.1 Pro Preview, Claude Fable 5, DeepSeek V4 Flash, Claude Sonnet 5, Mistral Medium 3.5

Questions people ask

What is the best AI for everyday maths and percentages in 2026?

Claude Opus 4.8 leads our tested ranking with 100/100 on our 12-task everyday-maths suite, in a dead heat with GPT-5.6 Terra (99). Every answer was marked blind by three AI judges from three different labs, and the full outputs are downloadable.

How is this ranking made?

Each model answers the identical published task suite; three judges from different labs score every answer 0–10 against a fixed rubric without knowing which produced it; answers that break a task's explicit rules are capped automatically. The score is the average, out of 100. No vendor pays for placement.

What happens when two models score the same?

They are separated by a fixed, published tie-break rather than by editorial choice: first the fewest machine-checked rule breaches, then the lowest measured cost per run. The order is therefore reproducible from the same data, and where the tie-break decides the top place the page says so.

How often does this page update?

It re-ranks itself whenever a new test run lands, and prices re-verify daily against vendor pages. The current ranking was last computed on 14 Aug 2026.

Head-to-heads in maths

Show all 20 tested pairs ▾

All comparisons →

More rankings ▾

Best AI for writing · Best AI chatbot for everyday use · Best AI for coding · Best free AI model · Best AI for spreadsheets and Excel · Best AI essay writer · Best AI for summarising documents · Best AI for extracting data from text · Best AI for writing emails · Best AI for customer service replies · Best AI for revision and study notes · Best AI for vibe coding · Best AI for making flashcards · Best AI for social media posts · Best AI for job applications and cover letters · Best AI for presentations · Best AI for writing your CV · Best AI for research skills · Best AI for creative writing · Best AI for translation · Best AI for travel planning · Best AI for emotional support · Best AI for everyday legal questions · Best AI for health questions · Best AI for writing a book · Best AI for humanising AI text · Best AI research agent · Best AI for code review · Best AI for workflow automation · Best AI for bookkeeping · Best AI for HR and employment questions · Best AI for meeting notes · Best AI for property and lettings · Best-value AI model API · every model we track · every tool