How this works
We test AI so you don't have to
Think Rotten Tomatoes, for AI. Every tool here has actually sat an exam — and you can read every answer it gave, and what the judges said about it.
30 tools · 21 models · 19 head-to-head battles · 654 scored test runs (every model×suite sitting, re-runs included) · last checked 11 Sept 2026
Everything sits the same exam
Each tool and model gets the identical set of real tasks — write a tricky work email, fix this bug, summarise these messy notes. Same questions, same conditions, so the scores mean something next to each other. The task suites are published.
Three rival AIs mark it blind
Every answer is scored by three judges from three different companies. None of them knows which tool wrote the answer, and none is ever from the same company as the thing it is marking. Break a checkable rule set by the task — a word limit, a banned phrase — and the harness itself catches it and caps the score in code; judges handle what needs reading, and we take the median of the three.
The public gets a vote too
Beside our tested score we show what actual users think: app-store ratings from millions of reviews, and opinions classified from 4,918 public posts across Reddit, YouTube, Hacker News and GitHub. Two scores, never averaged into one.
You can check every word of it
Each score links its receipts — the prompts, the raw answers, the judges’ verdicts. Each price links the vendor page we read it from, with the date. We hold 1,147 such facts, drawn from 172 source pages — all of it published beside the score.
What we will never do
- Take money for a ranking. No vendor pays for placement, position or inclusion. There is no affiliate link deciding an order.
- Score something we did not test. A tool with no test run shows no score. An empty page is better than an invented number.
- Call a tie a win. When two tools are too close to separate, the page says so and shows both.
- Hide a mistake. We found a scoring rule that had never actually been enforced, fixed it, re-ran everything, and left the old numbers visible.
Questions people ask
What is AI Intelligence?
A site that tests AI tools and models rather than listing them. 30 tools and 21 models sit identical task suites, every answer is marked blind by three AI judges from three different companies, and the score plus the full evidence is published.
How is it different from other 'best AI tool' pages?
Most are affiliate lists where nothing was tested, or benchmark sites written for researchers. Here every claim traces to evidence you can open: the task, the answer, the judge's own words, the vendor page a price came from and the date we read it.
How often does the site update?
A scan runs every hour for new models and price changes; sentiment and prices refresh each morning and the site rebuilds itself; a battle proposer picks its own next matchup daily. New releases trigger tests automatically, and rankings re-sort themselves when a result lands.
Why should I trust the scores?
A judge never marks its own company's model, marking is blind, and an answer that breaks a checkable rule — a word limit, a banned phrase — is caught by the test harness itself and capped in code. Our score sits beside what real users say, so you never have to take only our word. We publish our own mistakes too.
Does anyone pay to appear here?
No. No vendor pays for placement, ranking or inclusion, and there are no affiliate links deciding an order.
Start with a ranking — best everyday chatbot, best for writing, best free — or read the full testing protocol.