How we test
The short version: we run both contestants through the same published task suite via the same API harness, measure everything, and have three AI judges from three different labs score every pair blind, each in both presentation orders. The majority decides; a verdict that flips when the order is swapped does not count. We publish the raw outputs. Every number on this site traces to a citation with a source URL and a retrieval date. If a number can't be cited, it doesn't render.
Which tasks a verdict is scored on
A verdict is only about the thing it claims to be about. A skill category — everyday chat, writing, coding — is scored only on that skill's tasks (18 per skill, published in full below).
We got this wrong until 7 August 2026, and the correction is worth stating plainly. Every battle was scored on all 30 tasks of the general suite regardless of its category, so a page headed “Writing” was decided mostly by coding, reasoning and extraction. On re-running that battle on writing tasks alone the result moved materially. The old numbers were not wrong arithmetic; they were the wrong question. The corrected runs replaced the originals; the superseded scorelines are not retained on the site, which is a gap we are naming rather than papering over.
Two categories keep the general suite on purpose: best free and best-value API. There the question genuinely is overall capability, so a broad suite is the right test rather than a narrow one. Categories we cannot yet test honestly — research needs web-grounded models and meetings needs transcripts — say “in testing” instead of borrowing tasks that would not measure them. Image generation is tested on its own separate image suite with a vision judge, never on the text tasks.
The per-skill task suites — suite-2026-08-skills (54 tasks, published in full)
Per-skill task sets. A battle in a SKILL category (everyday chat, writing, coding) is scored only on that category's tasks — scoring a 'writing' verdict on coding and extraction tasks was never valid. Cross-cutting categories (best free, best-value API) still use the general suite in suite-2026-07.json, because there the question genuinely is overall capability. Tasks are written to have an objectively wrong answer wherever possible — a checkable constraint, a number that is right or wrong, a named edge case — because a task with no failure mode produces a tie, and a tie is not evidence.
Everyday chat & questions18 tasks — read every prompt
ec1 · Explain compound interest
Explain compound interest to a 15-year-old in no more than 120 words. Include one worked example with real numbers. Do not use the words 'exponential' or 'snowball'.
ec2 · Phone contract maths
Which is cheaper over three years: (a) a £35/month phone contract with a free handset, or (b) buying the handset outright for £300 plus a £12/month SIM-only plan? Show the arithmetic for both and state the winner and the difference.
ec3 · Recipe scaling
This carbonara serves 4: 320g spaghetti, 2 eggs, 1 egg yolk, 100g pancetta, 50g pecorino. Rescale it for 7 people. Give the new quantities, round sensibly for things you cannot buy in fractions, and say which one you rounded and why.
ec4 · Wifi troubleshooting
My laptop will not connect to my home wifi but my phone connects fine. Give me a numbered troubleshooting list, most likely cause first, maximum 8 steps. Each step must be an action I can actually take, not 'check your settings'.
ec5 · Cancel an appointment
Write a message cancelling a dentist appointment two hours beforehand. Apologise, give a brief reason, and ask to rebook next week. Maximum 60 words. Do not grovel.
ec6 · Virus vs bacteria
Explain the difference between a virus and a bacterium, and why antibiotics work on one and not the other. Maximum 100 words. Do not use any military or war analogy.
ec7 · Cook from what's in
I have chicken thighs, rice, one lemon, garlic and spinach, plus salt, pepper and oil. Give me one dinner recipe using only those. Include timings and temperatures. Do not add an ingredient I did not list.
ec8 · Offside rule
Explain the football offside rule to someone who has never watched a match, in under 80 words. Include the one thing people most often get wrong about it.
ec9 · Packing list
Give me a packing list for a weekend in the Scottish Highlands in October. Exactly five items, no more and no fewer, each with a one-line reason. Assume I already have clothes and a toothbrush.
ec10 · Child's question
My eight-year-old asks why the sky is blue. Answer the way I should say it to her: under 70 words, scientifically accurate, no 'because of the atmosphere' hand-waving.
ec11 · Landlord repair request
Draft a message to my landlord reporting a boiler that has stopped producing hot water, requesting urgent repair. Firm, not rude, and it should make clear this is a heating and hot water issue. Maximum 80 words.
ec12 · Rank by footprint
Rank these five foods by greenhouse gas emissions per kilogram of product, highest first: beef, cheese, chicken, tofu, lentils. Give an approximate figure in kg CO2e per kg for each and name the type of source that figure comes from.
ec13 · Couch to 5k
I cannot run at all and want to run 5k in eight weeks. Give me a week-by-week plan: exactly eight lines, one per week, each stating what I actually do that week. No preamble and no closing paragraph.
ec14 · Spot the scam
I got a text saying 'HMRC: you are due a refund of £284.50. Claim within 24 hours at hmrc-refund-claim.co.uk'. Tell me whether this is a scam, give me the three specific signals in this message that decide it, and tell me what to do next.
ec15 · Compare two decisions
I can either take a £3,000 pay rise or four extra days of annual leave. I earn £45,000 in England and work a standard 5-day week. Work out which is worth more in cash terms, showing your reasoning, and name one non-financial factor that should also count.
ec16 · Itinerary with constraints
Plan three days in Rome for a couple who hate queueing and love food. Maximum 200 words, one line per activity, morning/afternoon/evening for each day. Do not recommend anything that normally requires standing in a long line without saying how to avoid it.
ec17 · Explain a bill
My electricity bill shows: standing charge 60.10p/day, unit rate 24.50p/kWh, usage 320 kWh over 31 days. Work out the bill before VAT, then with 5% VAT. Show each step.
ec18 · Difficult message
A friend has asked to borrow £2,000 and I do not want to lend it. Write what I should send. Warm, clear, no false excuses, does not leave the door open. Maximum 90 words.
Writing18 tasks — read every prompt
w1 · Cold email
Write a cold email (maximum 120 words) from the owner of a 6-person UK accounting firm to a local brewery, offering a free first-month bookkeeping trial. Friendly but professional, one clear call to action, no buzzwords.
w2 · Product description
Write a 60-80 word product description for a stainless-steel insulated water bottle (750ml, keeps drinks cold 24h and hot 12h, leakproof, fits car cup holders). Audience: gym-goers. Do not use the phrase 'stay hydrated' or the word 'sleek'.
w3 · Summarise messy notes
Turn these meeting notes into a clean five-bullet summary a director could read in 20 seconds: 'ok so launch slipped again, marc says supplier issue?? maybe 3 wks. sarah wants to push the email campaign anyway but tom thinks we look stupid if we promo something not shippable. budget is fine. legal still havent signed the claims doc. next check in tues.' Exactly five bullets.
w4 · Tone rewrite
Rewrite this so it is warm, takes responsibility, and keeps every fact identical, in under 90 words: 'Your order was delayed because of a system error on our side. The refund has been processed and will arrive in 3-5 days. We cannot expedite it.'
w5 · Headline set
Write five headlines for a blog post about small UK retailers switching from Shopify to WooCommerce. Each must be under 60 characters. Number them. No colons and no questions.
w6 · Cut by half
Cut this to half its length without losing any factual content: 'We are pleased to be able to announce that, following a period of extensive consultation with our valued customers and partners across the region, we have taken the decision to extend our opening hours at the Brighton branch. From Monday 6th October, the branch will be open from 8am until 8pm on weekdays, and from 9am until 5pm on Saturdays. We very much hope that these extended hours will make it easier for our customers to visit us at a time that suits them.' State the original and new word counts.
w7 · Bad news email
Write an email telling a client their project will be two weeks late because we underestimated the integration work. Own it, no blame-shifting, offer one concrete mitigation, keep it under 130 words. Do not use the word 'unfortunately'.
w8 · Job advert
Write a job advert for a part-time bookkeeper at a 12-person UK design agency, 20 hours a week, hybrid, £32k pro rata. Maximum 180 words. Must include the salary and the hours. No 'rockstar', 'ninja' or 'family'.
w9 · Structured explainer
Explain to a non-technical small business owner what a payment chargeback is, why it happens, and what they should do when they get one. Use exactly three subheadings. Maximum 220 words total.
w10 · Voice match
Here is a brand's voice: short sentences, dry humour, never exclamation marks, addresses the reader as 'you', British spelling. Write a 70-word homepage intro in that voice for a company that repairs vintage watches.
w11 · Reply to a bad review
Write a public reply to this 2-star review: 'Food was fine but we waited 50 minutes for mains on a Tuesday with 6 tables occupied. Nobody said anything until I asked.' Acknowledge the specific failure, do not offer a generic apology, invite them back once, under 80 words.
w12 · Turn features into benefits
Rewrite these three features as benefits for a small e-commerce owner, one sentence each, no more than 20 words each: '256-bit encryption', 'REST API with webhooks', '99.95% uptime SLA'.
w13 · Constrained abstract
Summarise the following in exactly 40 words, no more, no fewer: 'A study of 1,240 UK small businesses found that those adopting automated invoicing reduced late payments by 23% on average within six months, but that firms with fewer than five employees saw no significant change, largely because their invoice volume was too low for the effect to register.' State the word count at the end.
w14 · Two audiences
Explain the same product update — 'we now support multi-currency invoicing' — twice. First for an existing customer in one sentence. Then for a finance director evaluating us, in three sentences. Label them A and B.
w15 · Remove the fluff
Rewrite this so it contains no marketing filler and only checkable statements: 'Our revolutionary AI-powered platform leverages cutting-edge machine learning to deliver unparalleled insights that transform how forward-thinking businesses unlock growth at scale.' If a claim cannot be made checkable, drop it and say what you dropped.
w16 · Sequence a launch
Write the subject lines and one-sentence bodies for a three-email launch sequence for a £49 online course on Amazon PPC. Emails go out on days 1, 3 and 6. Each subject line under 45 characters. Label the day on each.
w17 · Write to a deadline word count
Write a LinkedIn post about why most A/B tests on small e-commerce sites never reach significance. Between 90 and 110 words. No hashtags. No 'thoughts?' at the end. Open with a claim, not a question.
w18 · Faithful compression
Compress this into three bullets, preserving every number exactly: 'Q3 revenue was £412,000, up 8% year on year. Gross margin fell from 61% to 57% because of increased shipping costs. Headcount rose from 14 to 17, and we opened the Manchester office in August, which contributed £18,000 of the quarter's revenue.'
Coding18 tasks — read every prompt
c1 · Duration parser
Write a Python function parse_duration(s) that converts strings like '1h30m', '45s', '2h', '90m', '1h2m3s' into total seconds. Raise ValueError on anything malformed. Include three assert-based tests, one of which covers a malformed input.
c2 · Find the bug
This is meant to return the average of the positive numbers but returns the wrong value. Identify the bug, explain it in one sentence, and give the corrected function. function avgPositive(xs) { let sum = 0, n = 0; for (const x of xs) { if (x > 0) sum += x; n++; } return sum / n; }
c3 · SQL without window functions
Given tables users(id, email) and orders(id, user_id, created_at, total), write SQL returning the email and order count of every user with more than 3 orders in the last 30 days, most orders first. Do not use window functions. Target Postgres.
c4 · Infinite useEffect
Explain precisely why this React effect loops forever, then give the fixed version. const [items, setItems] = useState([]); useEffect(() => { fetch('/api/items').then(r => r.json()).then(setItems); }, [items]);
c5 · Typed debounce
Write a debounce function in TypeScript that preserves the argument types of the wrapped function, returns a function with a .cancel() method, and does not use 'any'. Explain in one sentence why the naive generic signature loses type information.
c6 · Leftmost binary search
Implement binary search that returns the index of the FIRST occurrence of a target in a sorted array with duplicates, or -1. Give the code and state the complexity. Include the test case that distinguishes it from an ordinary binary search.
c7 · Security review
Review this Express handler and list every security problem you find, most severe first, each with the fix. app.get('/file', (req, res) => { const p = req.query.name; db.query(`SELECT * FROM files WHERE name = '${p}'`, (e, rows) => { res.sendFile(__dirname + '/uploads/' + p); }); });
c8 · Safe migration
Write the Postgres migration to add a NOT NULL column 'status' with default 'pending' to an orders table with 40 million rows, without taking a long exclusive lock. Give the steps in order and say which step is the dangerous one and why.
c9 · Fix the code not the test
This test fails. Fix the implementation, not the test. // impl export const slugify = (s) => s.toLowerCase().replace(/ /g, '-'); // test expect(slugify(' Hello World! ')).toBe('hello-world');
c10 · Race condition
Identify the race condition in this code, explain what interleaving causes it, and fix it. let cache = null; async function getConfig() { if (cache) return cache; const r = await fetch('/config'); cache = await r.json(); return cache; }
c11 · Retry with backoff
Write an async retry wrapper in TypeScript: exponential backoff with jitter, a maximum attempt count, and it must NOT retry on 4xx responses other than 429. Maximum 30 lines. State what happens on the final failure.
c12 · Recursive type
Write a TypeScript type DeepPartial<T> that makes every nested property optional, and explain in one sentence how it must handle arrays differently from plain objects.
c13 · Bash one-liner
Give me a single shell command that finds the ten largest files under the current directory recursively and prints them human-readable, largest first. It must handle filenames containing spaces. Explain each part briefly.
c14 · Explain and cost
Explain what this does and give its time and space complexity, then rewrite it to be O(n). def has_dup(xs): for i in range(len(xs)): for j in range(i+1, len(xs)): if xs[i] == xs[j]: return True return False
c15 · Regex with limits
Write a regex that validates UK postcodes, give a one-line explanation of each part, and then state two valid UK postcodes your regex would reject or two invalid ones it would accept. Do not claim it is perfect.
c16 · Callback to async
Refactor this to async/await with correct error propagation. Errors must not be swallowed. getUser(id, (e, user) => { if (e) return cb(e); getOrders(user.id, (e2, orders) => { if (e2) return cb(e2); getTotals(orders, (e3, totals) => cb(e3, totals)); }); });
c17 · Diagnose from a trace
Given this Node stack trace, state the most likely root cause and the first thing you would check: TypeError: Cannot read properties of undefined (reading 'map') at renderRows (/app/src/table.js:42:19) at Table (/app/src/table.js:12:5) at renderWithHooks (/app/node_modules/react-dom/cjs/react-dom.development.js:16305:18) The component works in dev and fails only on the production build's first paint.
c18 · Idempotency
Design an idempotent POST /payments endpoint so a client retry cannot charge twice. Describe the key, where it is stored, what happens on a concurrent duplicate, and what you return the second time. Maximum 200 words. Name the failure mode your design still has.
The general task suite — suite-2026-07 (30 tasks, used by the cross-cutting categories: best free, best-value API)
v1 task suite. 6 suites x 5 tasks. Prompts are original, written for this suite, refreshed quarterly to resist contamination. Speed & cost is measured on every call rather than prompted separately; its 5 tasks below stress long/short generation so latency and token economics are visible.
Writing5 tasks — read every prompt
w1 · Cold email
Write a cold email (max 120 words) from the owner of a 6-person UK accounting firm to a local brewery, offering a free first-month bookkeeping trial. Friendly but professional, one clear call to action, no buzzwords.
w2 · Product description
Write a 60-80 word product description for a stainless-steel insulated water bottle (750ml, keeps drinks cold 24h/hot 12h, leakproof, fits car cup holders). Target audience: gym-goers. Avoid cliches like 'stay hydrated in style'.
w3 · Summarise messy notes
Turn these messy meeting notes into a clean 5-bullet summary a director could read in 20 seconds: 'ok so launch slipped again, marc says supplier issue?? maybe 3 wks. sarah wants to push the email campaign anyway but tom thinks wait. budget - we're 8k over, mostly the packaging redo. Q: do we tell retail partners now or after new date confirmed. also NEED to hire the warehouse temp before august rush. next mtg tues.'
w4 · Tone rewrite
Rewrite this complaint reply so it is warm, takes responsibility, and keeps the same facts, in under 90 words: 'Your order was delayed because of a system error on our side. The refund has been processed and will arrive in 3-5 days. We cannot offer further compensation. Let us know if you have questions.'
w5 · Headlines
Write 5 headline options (max 9 words each) for a blog post about how small restaurants can use AI to reduce food waste. Mix: 2 practical, 2 curiosity-driven, 1 with a number.
Coding5 tasks — read every prompt
c1 · Bug fix
This JavaScript function should return the median of a numeric array but gives wrong answers for even-length arrays and mutates the input. Fix both issues, return only the corrected function with a one-line explanation: function median(arr) { arr.sort(); const mid = Math.floor(arr.length / 2); return arr[mid]; }
c2 · Small feature
Write a Python function `chunk_invoices(invoices, max_total)` that takes a list of dicts like {'id': 'A1', 'amount': 120.5} and groups them into batches where each batch's summed amount does not exceed max_total. A single invoice larger than max_total goes in its own batch. Preserve order. Include 3 test cases using assert.
c3 · Explain code
Explain to a junior developer, in under 150 words, what this code does and one risk of using it: const cache = new Map(); function memo(fn) { return (...args) => { const k = JSON.stringify(args); if (!cache.has(k)) cache.set(k, fn(...args)); return cache.get(k); }; }
c4 · SQL query
Given tables orders(id, customer_id, created_at, total) and customers(id, name, country), write a single SQL query returning each country's top 3 customers by lifetime spend in 2025, with columns country, name, total_spend, rank. Use a window function. Standard PostgreSQL.
c5 · Regex
Write a regex that matches UK postcodes like 'SW1A 1AA', 'M1 1AE', 'B33 8TH' (allow lowercase and optional space), and a one-line JavaScript example using it to validate a form field. Briefly note one edge case your regex does NOT handle.
Reasoning5 tasks — read every prompt
r1 · Multi-step logic
A bakery sells loaves at £3.20. Ingredients cost £1.10/loaf, labour £0.90/loaf, fixed costs £480/week. They sell 450 loaves/week. A supplier offers ingredients at £0.85/loaf if they commit to 600 loaves/week of ingredients (unused ingredients are wasted). Current unsold rate is zero, and they could raise output to at most 520 loaves/week with £60/week extra labour cost overall. Should they take the deal? Show the profit calculation for both options and give a clear recommendation.
r2 · Planning
Plan the launch week for a small online course (already recorded). Resources: one founder, a part-time VA (10h), email list of 2,000, £300 ad budget. Produce a 7-day plan, one line per day, each line naming the owner. Flag the single riskiest dependency.
r3 · Trade-off analysis
A 12-person agency must choose: (A) hire a mid-level developer at £55k, or (B) contract overflow work to freelancers at roughly £400/day, expected 60 days/year. Give a recommendation in under 200 words covering cost, flexibility, quality risk, and one non-obvious factor.
r4 · Maths word problem
A tank holds 2,400 litres. Pump A fills at 40 L/min. Pump B drains at 25 L/min. A runs from 09:00. B accidentally switches on at 09:20. At 10:00 B is switched off. At what time is the tank full? Show your working.
r5 · Catch the flaw
Find the flaw in this argument and explain it in under 100 words: 'Our churn dropped from 6% to 3% after we introduced the loyalty programme in March, so the programme cut churn in half. We should double the loyalty budget to cut churn to 1.5%.'
Extraction5 tasks — read every prompt
e1 · Messy to JSON
Extract to JSON with keys name, company, phone, email, intent: 'hiya - jules moreau here from Atelier Nord (the lighting people). best number is 07700 900 457, or jules@ateliernord.fr. we're after a quote for the workshop series u mentioned, ideally b4 sept' Return only valid JSON.
e2 · Table from prose
Turn this into a markdown table with columns Product, Q1, Q2, Change: 'The Alpha line did 14,200 units in Q1 and 15,800 in Q2. Beta slipped from 9,100 to 7,400. The new Gamma launched mid-Q2 with 3,300 units (no Q1 sales). Delta held flat at 5,000 both quarters.' Include a Change column as a percentage to one decimal (write n/a where undefined).
e3 · Contact details
List every person mentioned below with their role and organisation, one line each in the format Name - Role - Org. If a field is unknown write unknown: 'Following the review, Dr Amara Okafor (Chief Data Officer, Halcyon Health) will hand over to Ben Tran, who joins as interim CDO from Meridian Labs. Legal sign-off sits with Priya Shah at Halcyon; the external audit remains with T. Lindqvist.'
e4 · Key facts
Extract exactly 5 key facts as bullets (each under 15 words) from: 'The council approved the harbour redevelopment on 12 June by 7 votes to 2. Phase one, costing £4.6m, begins in October and includes the fish market roof. The marina expansion (phase two) is unfunded and depends on a levelling-up grant decision expected in January. Local traders get first refusal on the 14 new retail units. Construction traffic will be routed via Dock Road, avoiding the high street.'
e5 · Schema compliance
Convert to JSON matching exactly this schema: {"event": string, "date": "YYYY-MM-DD", "venue": string, "capacity": number, "speakers": [{"name": string, "topic": string}], "sold_out": boolean} 'AI in Retail Summit happens March 14th 2026 at the Baltic Exchange (holds 320). Talks: Nadia Rossi on demand forecasting, Kwame Mensah on dynamic pricing. Tickets still available.' Return only the JSON.
Instruction-following5 tasks — read every prompt
i1 · Format constraints
Describe how photosynthesis works in exactly 3 sentences. The first sentence must start with 'Plants', the second with 'Inside', the third with 'Finally'. No sentence may exceed 20 words.
i2 · Word limit
Explain the difference between a lease and a licence for UK commercial property in exactly 50 words. Count carefully - exactly 50.
i3 · Forbidden words
Explain what makes coffee taste bitter, in 80-120 words, WITHOUT using any of these words: bitter, bean, roast, caffeine, acid. Do not use hyphenated or partial variants of them either.
i4 · Multi-constraint
Write a LinkedIn post about hiring your first apprentice. Constraints: 90-110 words, exactly one emoji at the very end, exactly 3 paragraphs, second paragraph must be a single question, include the phrase 'small wins compound', no hashtags.
i5 · Exact structure
Produce a checklist for closing a small retail shop at night with EXACTLY this structure: a title line, then 3 sections headed 'Front of house', 'Back office', 'Security', each containing exactly 4 numbered items, each item 3-8 words. Nothing else before or after.
Speed & cost5 tasks — read every prompt
s1 · Quick answer (short)
In one sentence: what does VAT stand for and what is the current standard UK rate?
s2 · Instant classification
Classify each as SPAM or NOT SPAM, output only two-column lines: 'You've won a £500 voucher click here' / 'Your Tesco order 8812 is ready for collection' / 'URGENT your account will be closed verify now' / 'Minutes from Thursday's board meeting attached'.
s3 · Short translation
Translate to French, keeping the informal tone: 'Hey, we're running 10 minutes late - grab us a table if you get there first?'
s4 · Medium generation
Write a 150-word 'About us' section for a family-run garden centre in Yorkshire founded in 1987, now run by the founders' daughter, known for rare perennials and a small cafe.
s5 · Long generation
Write a detailed 400-word beginner's guide to setting up a home compost bin: location, what to add, what to avoid, how long it takes, and one common mistake.
The judge protocol
Each pair of outputs is scored by three judges from three different labs, none of which is ever a contestant. Each judge sees the task and two anonymised responses and scores each 0–10 against a fixed rubric: correctness (weight 3), instruction-following (3), usefulness to a busy non-technical professional (2), clarity (2).
The constraint cap, and how we found it was decorative. If a task sets a checkable constraint — a word count, a banned word, a required structure — a response that breaks it cannot score above 5. That rule was in the rubric from the first run and, until August 2026, nothing enforced it. An audit of the stored judge reasoning found judges writing “capping score at 5” in prose and recording 7 in the same object: the instruction was read, acknowledged, and not carried out. Because our headline statistic is a test on score margins, an uncapped breach moved published p-values directly. Judges now report only whether each response broke the constraint; the cap is applied in code, the winner is re-derived from the capped scores, and each battle page publishes how many times it fired — a gate nobody can see fire is decoration.
Position swapping: every judge scores every pair twice, with the order reversed the second time. A judge whose verdict flips when the order flips has not produced a verdict, and its vote becomes a tie. This kills the well-documented first-position bias.
Why three, and why different labs: a single judge scoring both orders gave us a Cohen's κ of 0.10 on one category — barely better than chance, which means most of the “ties” we were publishing were the judge contradicting itself rather than the models being level. Judges from the same lab share failure modes, so they cannot average each other out. The task goes to the majority of the panel; no majority is a genuine tie. If a judge errors or returns something unparseable it abstains and the remaining judges decide — it is never counted as a vote for a tie. Panel agreement is published on each battle page.
Measured, not asked: latency, token counts and per-call cost are recorded by the harness on every request — they are measurements, not model claims.
Human spot-checks: judge decisions are spot-checked by a human before a battle is considered verified for indexing. Suites are refreshed quarterly to resist benchmark contamination.
The audience score — “theirs”, next to “ours”
Our verdict answers one question: which tool won a published task suite. It cannot tell you what the people who use these things every day actually think, and treating one as evidence for the other is how comparison sites become worthless. So we measure both, separately, and never average them into a single number.
How the audience score is made. A research sweep collects public posts about the tool from Reddit, Hacker News, YouTube and GitHub over the last 90 days. Every post is then classified one at a time as positive, negative or neutral about that tool, with a one-line reason. The score is positive / (positive + negative) — the same shape as a Rotten Tomatoes audience score.
Neutral posts are excluded, not split. Release notes, tutorials, changelogs and job ads carry no opinion, so counting them as half a vote would pad the sample and flatten every score toward the middle. They are shown in the sample count and left out of the percentage.
The exact rule, so you can reproduce any label on this site. We compute a 95% Wilson score interval on the positive share. We say mostly liked only if the whole interval sits above 50%, mostly criticised only if it sits entirely below, and opinion is splitwhen it genuinely straddles 50%. If the interval is wider than 40 points we say not enough data to say and make no claim at all — a wide interval is an absence of evidence, not a finding of balance. The width limit is 40 points.
Every score carries its own sample size. Under 8 posts carrying a clear view we show the raw counts and no percentage at all — a percentage implies a precision that few posts cannot support. Both of these numbers are read from the same constants the site renders with, so this paragraph cannot drift from the code: an earlier version of it quoted two different thresholds in consecutive sentences, neither of which was the one being applied. If a source was rate-limited or only partly reachable during a sweep, the page says so: a quiet source and a source we could not read are not the same thing.
It never touches the verdict. Sentiment does not move a measured number, does not feed the confidence gate, and does not decide what gets published. When the tasks and the crowd disagree we show both and say so plainly — that gap is the most useful thing on the page, not a problem to reconcile away. A tool the internet loves that our tasks do not reward is a real finding.
Public posts skew negative, and the score inherits that. People write about a tool when it breaks, changes price, or annoys them; satisfied users mostly say nothing. So an audience score is not an absolute quality rating, and a tool sitting at 20% is not “80% bad”. Every tool carries the same bias, which is what makes the number useful: it is meaningful when you compare two things of the same kind, and meaningless read on its own. For the same reason we never compare a tool’s score against a model’s — the people posting about the Claude app and the people posting about the Claude Sonnet API are different crowds having different conversations.
Every classified post is published. Each tool page carries the full set behind “Every post behind this score”: the headline, the source, the engagement, how it was classified and the classifier’s one-line reason, each linked so you can open it and disagree with us. For a period this page promised that and the site showed only the single most-engaged quote; an auditor caught it, and publishing the set was the fix rather than softening the promise.
“Is it safe?” — how we answer it without guessing
A reader who was not in the industry read this whole site and told us it answered a question she had not asked — which tool wins a coding test — and never answered the three she had: what happens to the things she types in, how she would know if it was making something up, and whether she has to hand over a card. Every tool page now opens with those, and none of the answers are ours.
The model is never asked what it knows. Asking a language model “does this company train on your chats?” produces a fluent, confident, unverifiable answer, and a wrong one is worse than silence because someone will act on it. Instead we fetch the vendor’s own privacy policy and terms — finding the URL from their own site, never from memory — and ask only for the sentence in that text that answers the question.
Every quote is then checked character-for-character. The sentence the model returns must appear verbatim in the document we fetched, after normalising whitespace and quote marks. If it does not, we treat it as invented and write nothing for that tool. That is the whole safeguard, and it means the pipeline routinely produces no answer: a page with no safety panel means we could not verify one, and a page with a safety panel means every line in it came out of a document the vendor published, with a link to the page and the date we read it.
One line on that panel is ours, and it says so. “What should you never type into it” is advice, not a fact about a vendor — it follows from what their policy says about training and retention, and it is labelled as our advice on its face. Where we could not read a policy, that advice assumes the worst rather than reassuring you.
What this does not cover. We read what companies publish. We do not audit whether they do it, we do not test their security, and a policy can change the day after we read it — which is why the date we read it sits next to every quote. Enterprise and business plans routinely have different terms from the consumer ones quoted here.
The confidence gate — when we refuse to declare a winner
A scoreline alone is not a verdict. Before we frame a battle as a win it must pass seven checks computed from the run itself, with no human able to override the fact-checking side and no machine able to skip the human side.
The decisiveness test changed on 7 August 2026, and the reason is worth stating. We used to run a sign test on win counts: every task was collapsed to win, loss or tie, and the margin was thrown away. So a 3.7-point thrashing and a 0.2-point nothing counted identically. With 18 tasks and about half of them tying, the effective sample was 9 — and at 9 you need to win 8 of them to clear the bar. Below 6 decided tasks it is not reachable at all, however one-sided the result. Seven battles produced zero verdicts, and that was a property of the test rather than a finding about AI tools.
We now run a Wilcoxon signed-rank test on the per-task score margins — the graded 0–10 scores the judges already give, which we were measuring and discarding. Exact null distribution for samples up to 22, normal approximation with a continuity correction above that. The old sign test is still computed and published alongside it as a secondary check, so nothing that used to be reported has quietly vanished. On our existing battles this moved p-values from 0.42 to 0.09, and from 0.23 to 0.06 — still short of the bar, but for the first time measuring what we actually collected.
Judge reliability is now a gate criterion, not just a published number. We measured Cohen’s κ and printed it, but a battle could still be called decisive on an instrument the same page admitted was barely better than chance. A verdict now requires κ of at least 0.21 — below that the judges are not consistent enough for a p-value to mean anything.
The threshold is p < 0.05, two-sided. It is the same for every battle and every category, it was set before any of these runs, and we do not move it to make a result publishable.
The other five checks are unchanged: a fact-check-passed article exists, at least 20 machine-verified citations, the data-completeness gate passes live, the scoreline recomputes from the raw task results, and cross-position disagreement is scored as a tie rather than forced into a verdict.
The citation rule
Every rendered number on this site is backed by a claim row carrying: the value, a source URL, the source name, the retrieval timestamp, and a verification status (verified / self-reported / stale). Where we have no verified claim, we render an honest empty state instead. There is no override for this in the codebase.
Community Elo scores come from the LMArena leaderboard dataset (CC-BY-4.0, attribution given). API prices come from the OpenRouter models API. App ratings come from Apple's iTunes Lookup API. Subscription prices are read from vendors' public pricing pages and dated.
Funding & independence
No vendor pays for placement — ever. No sponsored rankings, no "verified vendor" badges, no pay-to-rank. Some outbound links may become affiliate links; verdicts are computed from test results before any link exists, and the protocol above is the only input to a verdict.
Corrections policy
If we got a number wrong, we fix the claim, mark the old value superseded, and note the correction in the page changelog and the public changelog. Corrections are never silent. Spotted something? Email corrections@aiintelligence.com.