Best AI for everyday legal questions
Scored 94/100 on our 12-task legal-questions suite — level on quality with GPT-5.6 Sol (94) — the top spot goes on the tie-break: cleanest rule-compliance, then lowest measured cost; 4 more sit a single point behind.
Claude Opus 4.8 is the engine inside Claude. Go to claude.ai ↗ — the free tier is fine to start. Paid plans start at $20/month (pro plan, vendor’s own price). The free tier is Claude’s, not a promise about this exact model — we haven’t verified which plan carries it. Not fussed about the last point or two? Any of the top 5 here will serve you well.
England-flavoured everyday law: the landlord who 'lets himself in', the unprotected deposit, the 8-month-old washing machine the retailer disowns, the scary private parking letter, and a letter before action that stays civil. Marked on plain English, the honest 'it depends what you signed', and knowing when the only right answer is a solicitor. None of this is legal advice — answers are judged against public guidance, and your situation may differ. One task asks for help working around a non-molestation order — the right answer is to refuse.
updated 16 Aug 2026 · tested by Robert Prime · re-ranks automatically when a new run lands
| # | Model | Our score |
|---|---|---|
| 1 | Claude Opus 4.8 | 94/100 |
| 2 | GPT-5.6 Sollatest | 94/100 |
| 3 | GPT-5.6 Terralatest | 93/100 |
| 4 | GPT-5.3-Codex | 93/100 |
| 5 | GPT-5.5 | 93/100 |
| 6 | Claude Fable 5 | 93/100 |
| 7 | GPT-5.6 Luna | 92/100 |
| 8 | Kimi K3 | 91/100 |
| 9 | Claude Sonnet 5 | 88/100 |
| 10 | GLM 5.2 | 86/100 |
| 11 | Grok 4.5 | 86/100 |
| 12 | Qwen3.7 Max | 85/100 |
| 13 | Gemini 3.1 Pro Preview | 85/100 |
| 14 | DeepSeek V4 Pro | 83/100 |
| 15 | Gemini 3.5 Flash | 82/100 |
| 16 | Claude Opus 4.6 | 82/100 |
| 17 | Gemini 3.1 Flash Lite | 80/100 |
| 18 | DeepSeek V4 Flash | 73/100 |
| 19 | Mistral Medium 3.5 | 72/100 |
“API cost” is what software developers pay to build on a model — ignore it if you just use the website. Each model answers each task once. Models level on score are ranked by a fixed tie-break — fewest machine-checked rule breaches, then lowest measured cost per run — so the order is deterministic and checkable, never arbitrary. Judge panels never include the contestant’s own lab, so panels differ slightly per model — small cross-model gaps can reflect panel severity, not quality.
Made by Anthropic. You use it inside Claude — nothing to install.
Strongest showing: “Eight-month-old washing machine” — scored 10/10 by the panel. Weakest: “The handwritten will” at 9/10.
“The response perfectly addresses all aspects of the prompt. It accurately explains UK consumer law, the shift in the burden of proof at 6 months, and provides a highly effective, polite, and firm draft email.”— google/gemini-3.1-pro-preview, judging blind · full receipts ↓
12 tasks · 16 Aug 2026 · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5 · API $5 in / $25 out per 1M tokens · full model page →
Made by OpenAI — their newest model. You use it inside ChatGPT — nothing to install.
Strongest showing: “The hedge war” — scored 10/10 by the panel. Weakest: “It depends what you signed” at 9/10.
“The response perfectly addresses all constraints and requirements. It accurately outlines the escalation ladder, explains the high hedge complaint process, details its capabilities and limitations, explains why court is a last resort, and c”— google/gemini-3.1-pro-preview, judging blind · full receipts ↓
12 tasks · 16 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $5 in / $30 out per 1M tokens · full model page →
Made by OpenAI — their newest model. You use it inside ChatGPT — nothing to install.
Strongest showing: “Translate the clause” — scored 10/10 by the panel. Weakest: “Landlord with a key” at 8/10.
“The response perfectly translates the legal jargon into plain English, accurately identifies the promise and the carve-out, provides a highly relevant question, and stays well under the 80-word limit.”— google/gemini-3.1-pro-preview, judging blind · full receipts ↓
12 tasks · 16 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $1 in / $6 out per 1M tokens · full model page →
See it for yourself
One task from the suite, and what the leader, the middle of the table and the bottom actually wrote. This is the task they disagreed on most — scores ranged from 5 to 10 out of 10.
'Can my employer make me work Boxing Day? Everyone online says no.' Explain why the internet answer is wrong to be so certain — there is no general UK right to bank holidays off; it turns on the contract's wording. Show two or three contract phrasings that produce different answers, tell them exactly what to look for in their own contract, and what to do if they cannot find a copy of it.
# Can Your Employer Make You Work Boxing Day? The confident "no" you're seeing online is wrong — or at least, dangerously oversimplified. ## Why the internet is wrong to be so certain There is **no general legal right in the UK to have bank holidays off work**. Bank holidays are not automatically…
Judge: The response perfectly addresses all parts of the prompt. It accurately explains UK employment law regarding bank holidays, provides three distinct contract phrasings, details what to look for, and gives clear steps if the contract is missi
The reason the internet is so confidently wrong about this is a common misunderstanding of UK employment law. Many online forums conflate the *cultural* expectation of having bank holidays off with a *legal* right. In the UK, there is no general statutory right to have bank holidays off. Under the…
Judge: Accurate, well-structured, covers legal basis, three contract phrasings, what to check, and fallback steps including custom and practice. Clear and useful for non-technical reader with appropriate disclaimer.
The internet’s blanket "no" is wrong because **there is no automatic legal right in the UK to have bank holidays (including Boxing Day) off work**. Whether you must work depends entirely on **your employment contract’s wording**—not statute. Here’s why the online advice is misleading, and how to che…
Judge: The prompt explicitly asked to show 'two or three' contract phrasings, but the response provided four. Otherwise, the content is accurate, clear, and highly useful.
Trimmed to fit. Every full answer, every judge’s score: Claude Opus 4.8, GLM 5.2, Mistral Medium 3.5
How this ranking is made
Every model answers the identical published task suite. Three AI judges from three different labs mark each answer blind against a fixed 0–10 rubric — none is ever from the maker’s own lab, and an answer that breaks a task’s explicit rules is capped in code, not by the judge’s goodwill. The score is the average, out of 100.
What this is not: an opinion piece, an affiliate list, or a benchmark we cannot show you. Every score links its raw outputs and judge verdicts. The full protocol · receipts: Claude Opus 4.8, GPT-5.6 Sol, GPT-5.6 Terra, GPT-5.3-Codex, GPT-5.5, Claude Fable 5, GPT-5.6 Luna, Kimi K3, Claude Sonnet 5, GLM 5.2, Grok 4.5, Qwen3.7 Max, Gemini 3.1 Pro Preview, DeepSeek V4 Pro, Gemini 3.5 Flash, Claude Opus 4.6, Gemini 3.1 Flash Lite, DeepSeek V4 Flash, Mistral Medium 3.5
Questions people ask
What is the best AI for everyday legal questions in 2026?
Claude Opus 4.8 leads our tested ranking with 94/100 on our 12-task legal-questions suite (12 tasks), in a dead heat with GPT-5.6 Sol (94). Every answer was marked blind by three AI judges from three different labs, and the full outputs are downloadable.
How is this ranking made?
Each model answers the identical published task suite; three judges from different labs score every answer 0–10 against a fixed rubric without knowing which produced it; answers that break a task's explicit rules are capped automatically. The score is the average, out of 100. No vendor pays for placement.
How often does this page update?
It re-ranks itself whenever a new test run lands, and prices re-verify daily against vendor pages. The current ranking was last computed on 16 Aug 2026.
Head-to-heads in legal questions
Show all 20 tested pairs ▾
More rankings ▾
Best AI for writing · Best AI chatbot for everyday use · Best AI for coding · Best free AI model · Best AI for spreadsheets and Excel · Best AI essay writer · Best AI for summarising documents · Best AI for extracting data from text · Best AI for writing emails · Best AI for everyday maths and percentages · Best AI for customer service replies · Best AI for revision and study notes · Best AI for vibe coding · Best AI for making flashcards · Best AI for social media posts · Best AI for job applications and cover letters · Best AI for presentations · Best AI for writing your CV · Best AI for research skills · Best AI for creative writing · Best AI for translation · Best AI for travel planning · Best AI for emotional support · Best AI for health questions · Best AI for writing a book · Best AI for humanising AI text · Best-value AI model API · every model we track · every tool