GLM 5.2 vs Kimi K3 for everyday legal questions

The verdict

Kimi K3 wins, 9186

Same 12 everyday legal questions tasks, marked blind by three rival labs. Kimi K3 took 6 tasks, GLM 5.2 took 1, 5 tied. Tested 16 Aug 2026.

Where they differed most

The task — Translate the clause

A decorator's contract contains: 'The Client shall indemnify and hold harmless the Contractor against all claims, losses, damages and expenses howsoever arising in connection with the Works, save to the extent caused by the Contractor's negligence.' In max 80 words of plain English: what this clause makes the customer promise, the one carve-out it contains, and the question to ask before signing.

You promise to pay for any claims, losses, or damages related to the decorator's work. 

The one carve-out is that the decorator remains responsible for anything caused by their own negligence. 

Before signing, ask: "Does my home insurance cover this broad financial promise?"
Kimi K310/10
You promise to cover the decorator against any claim, loss, damage or expense connected with the work — however caused, even if it's not your fault (e.g., a visitor injured on site).

The only carve-out: you don't pay to the extent the loss was caused by the decorator's own negligence.

Before signing, ask: "Can this be limited to losses caused by my own acts, omissions or instructions?"

Task by task

TaskGLM 5.2Kimi K3
Landlord with a key89
The unprotected deposit89
Eight-month-old washing machine99
The scary parking letter99
Translate the clause810
The hedge war89
Hours cut overnight99
The handwritten will98
£600 and a shrug78
It depends what you signed99
Letter before action1010
Refuse the order workaround910

Full receipts: GLM 5.2, Kimi K3 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra

Questions people ask

Which is better for everyday legal questions: GLM 5.2 or Kimi K3?

Kimi K3 — it scored 91/100 against 86/100 on our 12-task everyday legal questions suite, winning 6 tasks to 1 with 5 tied. Every answer was marked blind by three judges from three rival AI labs.

How was this tested?

Both models answered the identical published everyday legal questions tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.

More everyday legal questions head-to-heads: Claude Opus 4.8 vs Kimi K3 · Claude Opus 4.8 vs GLM 5.2 · GPT-5.6 Sol vs Kimi K3 · GLM 5.2 vs GPT-5.6 Sol · GPT-5.6 Terra vs Kimi K3 · GLM 5.2 vs GPT-5.6 Terra

Full ranking: Best AI for everyday legal questions · model pages: GLM 5.2, Kimi K3