Best for / Spreadsheets
Best AI for spreadsheets and Excel
Spreadsheet work is where a confident wrong answer costs real money, so this suite is built almost entirely from checkable tasks: write the formula, fix the broken one, normalise the messy column, spot the number that cannot be right. A judge can point at the answer and say whether it works — no taste involved. One task deliberately asks for something impossible, to see which models say so instead of inventing a formula.
updated 13 Aug 2026 · tested by Robert Prime · re-ranks automatically when a new run lands
Scored 98/100 on our 12-task spreadsheet suite — in a dead heat with GPT-5.6 Luna (98), so either is a fine choice.
| # | Model | Our score | Context | Free? | API $/1M in·out |
|---|---|---|---|---|---|
| 1 | GPT-5.3-Codex | 98/100 | 400k | — | $1.75 · $14 |
| 2 | GPT-5.6 Lunalatest | 98/100 | 1050k | — | $0.1 · $0.6 |
| 3 | GPT-5.6 Terralatest | 98/100 | 1050k | — | $1 · $6 |
| 4 | GPT-5.5 | 97/100 | 1050k | — | $5 · $30 |
| 5 | GPT-5.6 Sollatest | 97/100 | 1050k | — | $5 · $30 |
| 6 | GLM 5.2latest | 96/100 | 1049k | chat.z.ai | $0.5 · $1.98 |
| 7 | Claude Opus 4.8 | 94/100 | 1000k | — | $5 · $25 |
| 8 | Grok 4.5latest | 94/100 | 500k | grok.com | $2 · $6 |
| 9 | DeepSeek V4 Flashlatest | 93/100 | 1049k | — | $0.14 · $0.28 |
| 10 | DeepSeek V4 Prolatest | 93/100 | 1049k | — | $1.168 · $2.336 |
| 11 | Gemini 3.1 Pro Preview | 93/100 | 1049k | — | $2 · $12 |
| 12 | Mistral Medium 3.5latest | 93/100 | 262k | — | $1.5 · $7.5 |
| 13 | Gemini 3.5 Flash | 92/100 | 1049k | Google AI Studio | $1.5 · $9 |
| 14 | Kimi K3latest | 90/100 | 1049k | — | $3 · $15 |
| 15 | Claude Fable 5 | 89/100 | 1000k | — | $10 · $50 |
| 16 | Gemini 3.1 Flash Lite | 83/100 | 1049k | — | $0.25 · $1.5 |
| 17 | Claude Opus 4.6 | 82/100 | 1000k | — | $5 · $25 |
| 18 | Claude Sonnet 5latest | 78/100 | 1000k | — | $2 · $10 |
openai's model, holding about 400k tokens of context — reached through the API.
Strongest showing: “Excel formula from a description” — scored 10/10 by the panel. Weakest: “Fix a broken formula” at 8/10.
“Correct, single formula, no pivot/helper column, clear concise explanation matching constraints.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $1.75 in / $14 out per 1M tokens · full model page →
openai's newest model, holding about 1050k tokens of context — reached through the API.
Strongest showing: “Excel formula from a description” — scored 10/10 by the panel. Weakest: “Fix a broken formula” at 8/10.
“Formula correctly sums last full month using SUMIFS with EOMONTH boundaries, no helper column/pivot, concise explanation.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $0.1 in / $0.6 out per 1M tokens · full model page →
openai's newest model, holding about 1050k tokens of context — reached through the API.
Strongest showing: “Excel formula from a description” — scored 10/10 by the panel. Weakest: “Fix a broken formula” at 8/10.
“Formula correctly sums last full calendar month, no helper column/pivot, clear explanation.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $1 in / $6 out per 1M tokens · full model page →
openai's model, holding about 1050k tokens of context — reached through the API.
Strongest showing: “Excel formula from a description” — scored 10/10 by the panel. Weakest: “Fix a broken formula” at 7/10.
“Formula correctly sums last full calendar month using SUMIFS with EOMONTH boundaries, no pivot/helper column, concise explanation.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $5 in / $30 out per 1M tokens · full model page →
openai's newest model, holding about 1050k tokens of context — reached through the API.
Strongest showing: “Excel formula from a description” — scored 10/10 by the panel. Weakest: “Fix a broken formula” at 6/10.
“Correct formula for last full month, single cell, no pivot/helper column, clear concise explanation.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $5 in / $30 out per 1M tokens · full model page →
z-ai's newest model, holding about 1049k tokens of context — reached through the API or free via chat.z.ai.
Strongest showing: “Excel formula from a description” — scored 10/10 by the panel. Weakest: “Fix a broken formula” at 7/10.
“Formula is correct, no helper column/pivot, single formula, clear explanation. Minor stylistic nitpicks only.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $0.5 in / $1.98 out per 1M tokens · free via chat.z.ai ↗ (checked 11 Aug 2026) · full model page →
anthropic's model, holding about 1000k tokens of context — reached through the API.
Strongest showing: “Excel formula from a description” — scored 10/10 by the panel. Weakest: “Google Sheets QUERY” at 5/10. The judges flagged 4 instruction breaches across the run — those scores were capped automatically.
“The formula is perfectly accurate and efficient. The response follows all constraints, including providing exactly one sentence of explanation, and is highly useful and clear.”— google/gemini-3.1-pro-preview, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5 · API $5 in / $25 out per 1M tokens · full model page →
x-ai's newest model, holding about 500k tokens of context — reached through the API or free via grok.com.
Strongest showing: “Clean messy data” — scored 10/10 by the panel. Weakest: “Google Sheets QUERY” at 5/10. The judges flagged 2 instruction breaches across the run — those scores were capped automatically.
“Correct, concise formula handling all cases via LET; trims spaces, uppercases, inserts space before last 3 chars. Minor: no TRIM for leading/trailing spaces beyond internal ones, but SUBSTITUTE handles that.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $2 in / $6 out per 1M tokens · free via grok.com ↗ (checked 11 Aug 2026) · full model page →
deepseek's newest model, holding about 1049k tokens of context — reached through the API.
Strongest showing: “Clean messy data” — scored 10/10 by the panel. Weakest: “Refuse the impossible” at 7/10.
“Formula correctly trims spaces, removes internal spaces, uppercases, and inserts space before last 3 chars, matching standard UK postcode format for all examples.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $0.14 in / $0.28 out per 1M tokens · full model page →
deepseek's newest model, holding about 1049k tokens of context — reached through the API.
Strongest showing: “Excel formula from a description” — scored 10/10 by the panel. Weakest: “Fix a broken formula” at 6/10. The judges flagged 1 instruction breach across the run — those scores were capped automatically.
“Correct formula, no helper column/pivot, clear explanation, meets constraints concisely.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $1.168 in / $2.336 out per 1M tokens · full model page →
google's model, holding about 1049k tokens of context — reached through the API.
Strongest showing: “Excel formula from a description” — scored 10/10 by the panel. Weakest: “Fix a broken formula” at 6/10. The judges flagged 3 instruction breaches across the run — those scores were capped automatically.
“Correct formula, no helper column/pivot, clear explanation, meets constraints and conciseness well.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gpt-5.6-terra, grok-4.5 · API $2 in / $12 out per 1M tokens · full model page →
mistralai's newest model, holding about 262k tokens of context — reached through the API.
Strongest showing: “Choose the right chart” — scored 10/10 by the panel. Weakest: “Refuse the impossible” at 7/10.
“Correct, clear, meets word limit, addresses chart type, axes, and limitation concisely; well within instructions for non-technical audience.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $1.5 in / $7.5 out per 1M tokens · full model page →
google's model, holding about 1049k tokens of context — reached through the API or free via Google AI Studio.
Strongest showing: “Google Sheets QUERY” — scored 10/10 by the panel. Weakest: “Fix a broken formula” at 6/10.
“Correct formula logic, appropriate GROUP BY/ORDER BY, header handling. Minor concern: header label true for row1 count, but fine since data starts row1.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gpt-5.6-terra, grok-4.5 · API $1.5 in / $9 out per 1M tokens · free via Google AI Studio ↗ (checked 11 Aug 2026) · full model page →
moonshotai's newest model, holding about 1049k tokens of context — reached through the API.
Strongest showing: “Excel formula from a description” — scored 10/10 by the panel. Weakest: “Fix a broken formula” at 3/10. The judges flagged 1 instruction breach across the run — those scores were capped automatically.
“Correct SUMIFS formula properly handling last full month, no helper column/pivot, concise explanation as required.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $3 in / $15 out per 1M tokens · full model page →
anthropic's model, holding about 1000k tokens of context — reached through the API.
Strongest showing: “Excel formula from a description” — scored 10/10 by the panel. Weakest: “Explain a formula back” at 5/10. The judges flagged 6 instruction breaches across the run — those scores were capped automatically.
“The formula is perfectly accurate and efficient. The explanation is exactly one sentence as requested, and all constraints are strictly followed.”— google/gemini-3.1-pro-preview, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5 · API $10 in / $50 out per 1M tokens · full model page →
google's model, holding about 1049k tokens of context — reached through the API.
Strongest showing: “Pivot logic in words” — scored 10/10 by the panel. Weakest: “Clean messy data” at 5/10.
“Clear, accurate, non-technical explanation, exactly 3 bullets, each under 20 words, avoids banned terms. Very effective for target audience.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gpt-5.6-terra, grok-4.5 · API $0.25 in / $1.5 out per 1M tokens · full model page →
anthropic's model, holding about 1000k tokens of context — reached through the API.
Strongest showing: “Clean messy data” — scored 10/10 by the panel. Weakest: “Fix a broken formula” at 4/10. The judges flagged 6 instruction breaches across the run — those scores were capped automatically.
“The response provides a flawless, elegant Excel formula using LET to perfectly normalize the postcodes. The explanation is clear, concise, and highly useful for a non-technical professional.”— google/gemini-3.1-pro-preview, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5 · API $5 in / $25 out per 1M tokens · full model page →
anthropic's newest model, holding about 1000k tokens of context — reached through the API.
Strongest showing: “Spot the error in the numbers” — scored 10/10 by the panel. Weakest: “Explain a formula back” at 5/10. The judges flagged 8 instruction breaches across the run — those scores were capped automatically.
“The response perfectly follows all instructions, providing a clear, accurate, one-sentence explanation of the likely data entry error and the corrected figure without speculating on business causes.”— google/gemini-3.1-pro-preview, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5 · API $2 in / $10 out per 1M tokens · full model page →
See it for yourself
One task from the suite, and what the leader, the middle of the table and the bottom actually wrote. This is the task they disagreed on most — scores ranged from 5 to 10 out of 10.
In Google Sheets, tab 'Sales' has headers in row 1: Date, Rep, Region, Amount. Write one QUERY() formula returning total Amount per Rep for the North region only, sorted highest first. Return only the formula.
=QUERY(Sales!A1:D,"select B, sum(D) where C = 'North' group by B order by sum(D) desc label B 'Rep', sum(D) 'Total Amount'",1)
Judge: Correct QUERY formula, proper syntax, sorted descending, grouped by Rep; header row included correctly. Minor: no forbidden constraints violated.
```sql =QUERY(Sales!A:D, "SELECT B, SUM(D) WHERE C = 'North' GROUP BY B ORDER BY SUM(D) DESC LABEL SUM(D) ''", 1) ```
Judge: Formula is correct logically; minor issue: wrapped in ```sql code block instead of plain text, and header row inclusion could affect result count, but formula itself is accurate.
``` =QUERY(Sales!A1:D, "select Col2, sum(Col4) where Col3 = 'North' group by Col2 order by sum(Col4) desc label Col2 'Rep', sum(Col4) 'Amount'", 1) ```
Judge: Formula uses 'Col2' syntax which fails on standard ranges (requires 'B', 'C', etc. unless wrapped in an array). Also violates 'only the formula' constraint by including markdown backticks.
Trimmed to fit. Every full answer, every judge’s score: GPT-5.3-Codex, DeepSeek V4 Pro, Claude Sonnet 5
How this ranking is made
Every model answers the identical published task suite. Three AI judges from three different labs mark each answer blind against a fixed 0–10 rubric — none is ever from the maker’s own lab, and an answer that breaks a task’s explicit rules is capped in code, not by the judge’s goodwill. The score is the average, out of 100.
What this is not: an opinion piece, an affiliate list, or a benchmark we cannot show you. Every score links its raw outputs and judge verdicts. The full protocol · receipts: GPT-5.3-Codex, GPT-5.6 Luna, GPT-5.6 Terra, GPT-5.5, GPT-5.6 Sol, GLM 5.2, Claude Opus 4.8, Grok 4.5, DeepSeek V4 Flash, DeepSeek V4 Pro, Gemini 3.1 Pro Preview, Mistral Medium 3.5, Gemini 3.5 Flash, Kimi K3, Claude Fable 5, Gemini 3.1 Flash Lite, Claude Opus 4.6, Claude Sonnet 5
Questions people ask
What is the best AI for spreadsheets and Excel in 2026?
GPT-5.3-Codex leads our tested ranking with 98/100 on our 12-task spreadsheet suite (12 tasks), in a dead heat with GPT-5.6 Luna (98). Every answer was marked blind by three AI judges from three different labs, and the full outputs are downloadable.
How is this ranking made?
Each model answers the identical published task suite; three judges from different labs score every answer 0–10 against a fixed rubric without knowing which produced it; answers that break a task's explicit rules are capped automatically. The score is the average, out of 100. No vendor pays for placement.
How often does this page update?
It re-ranks itself whenever a new test run lands, and prices re-verify daily against vendor pages. The current ranking was last computed on 13 Aug 2026.
More rankings: Best AI for writing · Best AI chatbot for everyday use · Best AI for coding · Best free AI model · Best AI for writing essays · Best AI for summarising documents · Best AI for extracting data from text · Best-value AI model API · every model we track · every tool