GLM 5.2 vs Kimi K3 for everyday legal questions
Kimi K3 wins, 91–86
Same 12 everyday legal questions tasks, marked blind by three rival labs. Kimi K3 took 6 tasks, GLM 5.2 took 1, 5 tied. Tested 16 Aug 2026.
Where they differed most
A decorator's contract contains: 'The Client shall indemnify and hold harmless the Contractor against all claims, losses, damages and expenses howsoever arising in connection with the Works, save to the extent caused by the Contractor's negligence.' In max 80 words of plain English: what this clause makes the customer promise, the one carve-out it contains, and the question to ask before signing.
You promise to pay for any claims, losses, or damages related to the decorator's work. The one carve-out is that the decorator remains responsible for anything caused by their own negligence. Before signing, ask: "Does my home insurance cover this broad financial promise?"
You promise to cover the decorator against any claim, loss, damage or expense connected with the work — however caused, even if it's not your fault (e.g., a visitor injured on site). The only carve-out: you don't pay to the extent the loss was caused by the decorator's own negligence. Before signing, ask: "Can this be limited to losses caused by my own acts, omissions or instructions?"
Task by task
| Task | GLM 5.2 | Kimi K3 |
|---|---|---|
| Landlord with a key | 8 | 9 |
| The unprotected deposit | 8 | 9 |
| Eight-month-old washing machine | 9 | 9 |
| The scary parking letter | 9 | 9 |
| Translate the clause | 8 | 10 |
| The hedge war | 8 | 9 |
| Hours cut overnight | 9 | 9 |
| The handwritten will | 9 | 8 |
| £600 and a shrug | 7 | 8 |
| It depends what you signed | 9 | 9 |
| Letter before action | 10 | 10 |
| Refuse the order workaround | 9 | 10 |
Full receipts: GLM 5.2, Kimi K3 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra
Questions people ask
Which is better for everyday legal questions: GLM 5.2 or Kimi K3?
Kimi K3 — it scored 91/100 against 86/100 on our 12-task everyday legal questions suite, winning 6 tasks to 1 with 5 tied. Every answer was marked blind by three judges from three rival AI labs.
How was this tested?
Both models answered the identical published everyday legal questions tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.
More everyday legal questions head-to-heads: Claude Opus 4.8 vs Kimi K3 · Claude Opus 4.8 vs GLM 5.2 · GPT-5.6 Sol vs Kimi K3 · GLM 5.2 vs GPT-5.6 Sol · GPT-5.6 Terra vs Kimi K3 · GLM 5.2 vs GPT-5.6 Terra
Full ranking: Best AI for everyday legal questions · model pages: GLM 5.2, Kimi K3