Which AI model is best right now

Measured, not asserted: GPT-5.6 Sol leads on 95.5/100, averaged across 30 job suites we publish in full.

Updated 28 August 2026 · twelve tasks per suite · every answer marked blind by three rival labs

#ModelMeanJobs satStrongest / weakestAPI $/1M
1GPT-5.6 Sol95.5/10030 of 30customer service replies 100
multi-step research 89
$2·$10
2GPT-5.595.2/10030 of 30customer service replies 99
workflow automation 88
$5·$30
3GPT-5.3-Codex94.3/10030 of 30extracting data from text 100
travel planning 88
$1.75·$14
4GPT-5.6 Terra94.3/10030 of 30extracting data from text 100
essay writing 88
$2·$12
5GPT-5.6 Luna93.3/10030 of 30extracting data from text 98
bookkeeping 83
$0.2·$1.20
6Claude Opus 4.890.3/10030 of 30everyday maths 100
essay writing 79
$5·$25
7Claude Fable 590.2/10024 of 30refused 6emotional support 97
revision & study 80
$10·$50
8Grok 4.5free via grok.com89.9/10030 of 30extracting data from text 100
travel planning 83
$2·$6
9Qwen3.7 Max89.1/10030 of 30extracting data from text 100
multi-step research 78
$1.48·$4.42
10Kimi K387.9/10030 of 30extracting data from text 100
job applications 78
$3·$15
11DeepSeek V4 Pro87.7/10030 of 30extracting data from text 100
vibe coding 63
$0.87·$1.74
12GLM 5.2free via chat.z.ai87.7/10030 of 30everyday maths 97
travel planning 76
$1.19·$3.74
13Gemini 3.1 Pro Preview87.3/10030 of 30extracting data from text 100
bookkeeping 78
$2·$12
14Claude Sonnet 587.2/10030 of 30emotional support 98
extracting data from text 67
$2·$10
15Gemini 3.5 Flashfree via Google AI Studio85.6/10030 of 30extracting data from text 100
vibe coding 68
$1.50·$9
16Claude Opus 4.684.6/10030 of 30everyday maths 98
extracting data from text 59
$5·$25
17DeepSeek V4 Flash83.1/10030 of 30extracting data from text 96
travel planning 68
$0.0886·$0.1772
18Gemini 3.1 Flash Lite80.8/10030 of 30extracting data from text 100
bookkeeping 67
$0.25·$1.50
19Mistral Medium 3.578.3/10030 of 30spreadsheets & Excel 93
bookkeeping 65
$1.50·$7.50

The mean is the average of one score per job suite, each from twelve tasks marked blind by three rival labs, and each linked above to the board where the tasks and the models’ real answers are printed in full. It is deliberately not a benchmark score: it measures the jobs people described to us rather than exam questions, and every number behind it is published. Read the two columns beside it before you read the order. “Jobs sat” is coverage — anything under 20 of 30 is listed but not ranked. “Refused” counts suites a model declined outright, which is its own answer and not a test we skipped. And a model whose strongest and weakest jobs are twenty points apart is not really described by its average at all.

Or pick the job you actually need

The overall table is a starting point, not an answer — the winner changes by job, and the cheapest model beats the flagship on more of these than you would expect.

Compare any models side by side → · How we test →