Best free AI model
Scored 92/100 on our 30-task general suite — level on quality with Grok 4.5 (92) — the top spot goes on the tie-break: cleanest rule-compliance, then lowest measured cost. Free to use via chat.z.ai.
GLM 5.2 is the engine inside Z Chat. Go to chat.z.ai ↗ — free to use. Not fussed about the last point or two? Any of the top 2 here will serve you well.
"Free" is the most abused word in AI marketing, so this page has a rule: a model only appears here if we verified its £0 route on the vendor's own site — the page, the date, the link. Quality scores come from the same 30-task general suite as the paid models, so you can see exactly what paying would buy you.
updated 11 Aug 2026 · tested by Robert Prime · re-ranks automatically when a new run lands
- All 3 sat the identical 30-task suite — same tasks, same order, one attempt each.
- Every answer marked blind: judges are not told which entrant wrote it.
- Three judges per answer, each from a competing lab. A panel never includes the entrant's own lab.
- Scored 0–10 against a fixed rubric. Answers breaking a task's explicit rules are capped by machine, not by opinion.
- Ties broken by fewest rule breaches, then lowest measured cost — published, not editorial.
- Nobody pays for placement. Last computed 11 Aug 2026.
Every score below links to the raw file behind it: every task, the entrant’s real answer, and all three judges’ marks.
| # | Model | Our score |
|---|---|---|
| 1 | GLM 5.2latest | 92/100 |
| 2 | Grok 4.5latest | 92/100 |
| 3 | Gemini 3.5 Flash | 89/100 |
“API cost” is what software developers pay to build on a model — ignore it if you just use the website. Each model answers each task once. Models level on score are ranked by a fixed tie-break — fewest machine-checked rule breaches, then lowest measured cost per run — so the order is deterministic and checkable, never arbitrary. Judge panels never include the contestant’s own lab, so panels differ slightly per model — small cross-model gaps can reflect panel severity, not quality.
Made by Z.ai — their newest model. You use it inside Z Chat — nothing to install; free via chat.z.ai.
Strongest showing: “Medium generation” — scored 10/10 by the panel. Weakest: “Explain code” at 6/10. On 3 tasks a judge flagged a possible instruction breach — advisory only: the machine check did not confirm it and the panel outvoted it.
“The response flawlessly incorporates all required details into a warm, engaging narrative. It perfectly hits the 150-word constraint (counting hyphenated words as single words) and is highly suitable for a business website.”— google/gemini-3.1-pro-preview, judging blind · full receipts ↓
30 tasks · 11 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $0.966 in / $3.036 out per 1M tokens · free via chat.z.ai ↗ (checked 11 Aug 2026) · full model page →
Made by xAI — their newest model. You use it inside Grok — nothing to install; free via grok.com.
Strongest showing: “Contact details” — scored 10/10 by the panel. Weakest: “Forbidden words” at 7/10. On 3 tasks a judge flagged a possible instruction breach — advisory only: the machine check did not confirm it and the panel outvoted it.
“The response perfectly follows the instructions, accurately extracting all four individuals and formatting them exactly as requested, correctly using 'unknown' where the organisation is not specified.”— google/gemini-3.1-pro-preview, judging blind · full receipts ↓
30 tasks · 11 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $2 in / $6 out per 1M tokens · free via grok.com ↗ (checked 11 Aug 2026) · full model page →
Made by Google. You use it inside Gemini — nothing to install; free via Google AI Studio.
Strongest showing: “Long generation” — scored 8/10 by the panel. On 6 tasks a judge flagged a possible instruction breach — advisory only: the machine check did not confirm it and the panel outvoted it.
“Accurate, well-organized, covers all requested elements clearly; slightly over 400 words but close, minor length deviation not strictly checkable.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
30 tasks · 11 Aug 2026 · judges claude-sonnet-5, gpt-5.6-terra, grok-4.5 · API $1.5 in / $9 out per 1M tokens · free via Google AI Studio ↗ (checked 27 Aug 2026) · full model page →
See it for yourself
One task from the suite, and what the leader, the middle of the table and the bottom actually wrote. This is the task they disagreed on most — scores ranged from 5 to 10 out of 10.
Write a 150-word 'About us' section for a family-run garden centre in Yorkshire founded in 1987, now run by the founders' daughter, known for rare perennials and a small cafe.
Welcome to our Yorkshire garden centre, a family-run haven nestled in the heart of the countryside. Founded in 1987 by my parents, what began as a small polytunnel and a big dream has blossomed into a beloved local institution. Today, I am proud to carry on their legacy as the second generation to n…
Judge: The response flawlessly incorporates all required details into a warm, engaging narrative. It perfectly hits the 150-word constraint (counting hyphenated words as single words) and is highly suitable for a business website.
**About Us** Nestled in the heart of Yorkshire, our family-run garden centre has been a cherished local haven since 1987. Founded by passionate horticulturists with a love for the land, it began as a modest nursery dedicated to quality plants and personal service. Today, it thrives under the care o…
Judge: The response includes all required details and is well-written, but it is 156 words long, violating the strict 150-word constraint.
Nestled in the heart of the beautiful Yorkshire countryside, our independent, family-run garden centre has been nurturing a passion for plants since 1987. Originally founded by our parents, the nursery is now proudly run by their daughter, who continues the family legacy with the very same dedicatio…
Judge: Warm, accurate, and well-structured copy covering the family history, daughter’s leadership, rare perennials, and café. It is concise and useful, but contains 149 words rather than the required 150.
Trimmed to fit. Every full answer, every judge’s score: GLM 5.2, Grok 4.5, Gemini 3.5 Flash
How this ranking is made
Every model answers the identical published task suite. Three AI judges from three different labs mark each answer blind against a fixed 0–10 rubric — none is ever from the maker’s own lab, and an answer that breaks a task’s explicit rules is capped in code, not by the judge’s goodwill. The score is the average, out of 100.
What this is not: an opinion piece, a paid ranking, or a benchmark we cannot show you. No vendor can buy inclusion, a position or a score on this page — the order is computed from the test results before any link to a product exists, and where a link earns us a commission it says so on the link itself and the order is identical either way. Every score links its raw outputs and judge verdicts. The full protocol · How we make money · receipts: GLM 5.2, Grok 4.5, Gemini 3.5 Flash
Questions people ask
What is the best free AI model in 2026?
GLM 5.2 leads our tested ranking with 92/100 on our 30-task general suite, in a dead heat with Grok 4.5 (92). Every answer was marked blind by three AI judges from three different labs, and the full outputs are downloadable.
How is this ranking made?
Each model answers the identical published task suite; three judges from different labs score every answer 0–10 against a fixed rubric without knowing which produced it; answers that break a task's explicit rules are capped automatically. The score is the average, out of 100. No vendor pays for placement.
What happens when two models score the same?
They are separated by a fixed, published tie-break rather than by editorial choice: first the fewest machine-checked rule breaches, then the lowest measured cost per run. The order is therefore reproducible from the same data, and where the tie-break decides the top place the page says so.
How often does this page update?
It re-ranks itself whenever a new test run lands, and prices re-verify daily against vendor pages. The current ranking was last computed on 11 Aug 2026.
Are these models really free?
Each model on this page carries a verified £0 route: GLM 5.2 via chat.z.ai; Grok 4.5 via grok.com; Gemini 3.5 Flash via Google AI Studio. We confirmed each on the vendor's own site, with the date shown. Models whose free access we could not verify are not listed here.
Is GLM 5.2 free?
Yes — GLM 5.2 can be used at no cost via chat.z.ai, confirmed on 11 Aug 2026.
Head-to-head battles
Worth knowing: Gemini 3.5 Flash finished ahead of GLM 5.2 head-to-head despite ranking below it above; Gemini 3.5 Flash finished ahead of Grok 4.5 head-to-head despite ranking below it above. The ranking scores each one alone against a fixed rubric; a battle puts them on the same task and asks which answer is better. When a gap this small flips, the honest read is that they are hard to separate — not that one of the numbers is wrong.
More rankings ▾
Best AI for writing · Best AI chatbot for everyday use · Best AI for coding · Best AI for spreadsheets and Excel · Best AI essay writer · Best AI for summarising documents · Best AI for extracting data from text · Best AI for writing emails · Best AI for everyday maths and percentages · Best AI for customer service replies · Best AI for revision and study notes · Best AI for vibe coding · Best AI for making flashcards · Best AI for social media posts · Best AI for job applications and cover letters · Best AI for presentations · Best AI for writing your CV · Best AI for research skills · Best AI for creative writing · Best AI for translation · Best AI for travel planning · Best AI for emotional support · Best AI for everyday legal questions · Best AI for health questions · Best AI for writing a book · Best AI for humanising AI text · Best AI research agent · Best AI for code review · Best AI for workflow automation · Best AI for bookkeeping · Best AI for HR and employment questions · Best AI for meeting notes · Best AI for property and lettings · Best-value AI model API · every model we track · every tool