Which AI model is best right now
Measured, not asserted: GPT-5.6 Sol leads on 95.5/100, averaged across 30 job suites we publish in full.
Updated 28 August 2026 · twelve tasks per suite · every answer marked blind by three rival labs
The mean is the average of one score per job suite, each from twelve tasks marked blind by three rival labs, and each linked above to the board where the tasks and the models’ real answers are printed in full. It is deliberately not a benchmark score: it measures the jobs people described to us rather than exam questions, and every number behind it is published. Read the two columns beside it before you read the order. “Jobs sat” is coverage — anything under 20 of 30 is listed but not ranked. “Refused” counts suites a model declined outright, which is its own answer and not a test we skipped. And a model whose strongest and weakest jobs are twenty points apart is not really described by its average at all.
Or pick the job you actually need
The overall table is a starting point, not an answer — the winner changes by job, and the cheapest model beats the flagship on more of these than you would expect.