Take the data. It's yours.

Every score we publish, as a file. 592 rows covering 19 models across 33 jobs, plus every suite a model refused. Use it for anything, including commercially — the only condition is that you credit us.

Rebuilt 23 September 2026 · CC BY 4.0 · no sign-up, no key, no rate limit

Why it’s free

This site’s whole argument is that you should not take a ranking on trust — that the tasks, the answers and the judges’ marks should all be open to inspection. Keeping the numbers locked in an HTML table while saying that would be a slightly embarrassing position to hold. So here is the table.

There is a self-interested half too, and it is only fair to say it out loud: the licence asks for credit, credit on the web means a link, and links are the thing an independent testing site cannot buy without becoming the sort of site it exists to warn people about. If this data is useful to you, the trade is a fair one.

How to cite it

AI Intelligence (2026). Blind-judged AI model scores across 33 job suites. Retrieved from https://aiintelligence.com/data

Or simply: “tested by AI Intelligence” with a link. If you are writing about a single score, link the specific board instead — every row in the CSV carries the exact page and the raw receipt file it came from.

What each column means

ColumnWhat it is
score_out_of_100Mean of three blind judges' 0–10 marks across every task in the suite.
rule_breachesHow many machine-checked task rules the model broke. The first tie-break, and often more telling than the score.
judgesThe three models that marked it, each from a lab other than the one being marked.
tasksHow many tasks in the suite. Identical for every model in a given job.
run_cost_usdWhat the run actually cost us in API spend. Published so the cost of the evidence is checkable too.
measured_atWhen the run finished. Scores are re-run, so this is the date the number was true.
receiptsPath to the raw file: every task, the model's real answer, all three judges' marks.

Three things to know before you use it

Refusals are a separate file, not missing rows. One model declines six of these jobs outright. Dropping those would quietly flatter it against models that attempted everything, so they are exported with the provider’s own refusal reason and the date.

Judges are LLMs, and we say so. Three of them, from competing labs, marking blind against a fixed rubric, never including the lab being marked. That is a real protocol and it is also not a human expert panel. The receipts let you disagree with any individual mark, which is the point of publishing them.

A score describes one suite on one date. Models change under the same name. Every row is stamped, and re-runs replace rather than average, so the file always holds the most recent measurement rather than a blend of old and new.

How the testing works → · The rankings this data produces → · Methodology in full →