Which AI model is best right now

Measured, not asserted: GPT-5.6 Sol leads on 95.4/100, averaged across 31 job suites we publish in full.

Updated 9 September 2026 · twelve tasks per suite · every answer marked blind by three rival labs

#ModelMeanJobs satStrongest / weakestAPI $/1M
1GPT-5.6 Sol95.4/10031 of 31customer service replies 100
multi-step research 89
$2·$10
2GPT-5.595.1/10031 of 31customer service replies 99
workflow automation 88
$5·$30
3GPT-5.3-Codex94.2/10031 of 31extracting data from text 100
travel planning 88
$1.75·$14
4GPT-5.6 Terra94.1/10031 of 31extracting data from text 100
meeting notes 88
$2·$12
5GPT-5.6 Luna93.1/10031 of 31extracting data from text 98
bookkeeping 83
$0.2·$1.20
6Claude Opus 4.890.2/10031 of 31everyday maths 100
essay writing 79
$5·$25
7Grok 4.5free via grok.com90.1/10031 of 31extracting data from text 100
travel planning 83
$2·$6
8Claude Fable 589.7/10025 of 31refused 6emotional support 97
meeting notes 79
$10·$50
9Qwen3.7 Max89.1/10031 of 31extracting data from text 100
multi-step research 78
$1.48·$4.42
10Kimi K387.9/10031 of 31extracting data from text 100
job applications 78
$3·$15
11GLM 5.2free via chat.z.ai87.7/10031 of 31everyday maths 97
travel planning 76
$0.966·$3.04
12DeepSeek V4 Pro87.6/10031 of 31extracting data from text 100
vibe coding 63
$0.9553·$1.91
13Gemini 3.1 Pro Preview87.3/10031 of 31extracting data from text 100
bookkeeping 78
$2·$12
14Claude Sonnet 587.2/10031 of 31emotional support 98
extracting data from text 67
$2·$10
15Gemini 3.5 Flashfree via Google AI Studio85.6/10031 of 31extracting data from text 100
vibe coding 68
$1.50·$9
16Claude Opus 4.684.7/10031 of 31everyday maths 98
extracting data from text 59
$5·$25
17DeepSeek V4 Flash83.1/10031 of 31extracting data from text 96
travel planning 68
$0.0886·$0.1772
18Gemini 3.1 Flash Lite80.6/10031 of 31extracting data from text 100
bookkeeping 67
$0.25·$1.50
19Mistral Medium 3.578.5/10031 of 31spreadsheets & Excel 93
bookkeeping 65
$1.50·$7.50

The mean is the average of one score per job suite, each from twelve tasks marked blind by three rival labs, and each linked above to the board where the tasks and the models’ real answers are printed in full. It is deliberately not a benchmark score: it measures the jobs people described to us rather than exam questions, and every number behind it is published. Read the two columns beside it before you read the order. “Jobs sat” is coverage — anything under 20 of 31is listed but not ranked. “Refused” counts suites a model declined outright, which is its own answer and not a test we skipped. And a model whose strongest and weakest jobs are twenty points apart is not really described by its average at all.

Or pick the job you actually need

The overall table is a starting point, not an answer — the winner changes by job, and the cheapest model beats the flagship on more of these than you would expect.

Compare any models side by side → · How we test →