GLM 5.2 vs GPT-5.6 Terra for everyday legal questions
GPT-5.6 Terra wins, 93–86
Same 12 everyday legal questions tasks, marked blind by three rival labs. GPT-5.6 Terra took 7 tasks, GLM 5.2 took 0, 5 tied. Tested 16 Aug 2026.
Where they differed most
A decorator's contract contains: 'The Client shall indemnify and hold harmless the Contractor against all claims, losses, damages and expenses howsoever arising in connection with the Works, save to the extent caused by the Contractor's negligence.' In max 80 words of plain English: what this clause makes the customer promise, the one carve-out it contains, and the question to ask before signing.
You promise to pay for any claims, losses, or damages related to the decorator's work. The one carve-out is that the decorator remains responsible for anything caused by their own negligence. Before signing, ask: "Does my home insurance cover this broad financial promise?"
The customer promises to cover the decorator’s costs and liabilities for any claims, losses, damage or expenses connected with the work. The exception is where those costs were caused by the decorator’s own negligence. Before signing, ask: “Why am I taking responsibility for all work-related claims, including risks I may not control, and can this be limited to losses caused by me?”
Task by task
| Task | GLM 5.2 | GPT-5.6 Terra |
|---|---|---|
| Landlord with a key | 8 | 8 |
| The unprotected deposit | 8 | 9 |
| Eight-month-old washing machine | 9 | 10 |
| The scary parking letter | 9 | 9 |
| Translate the clause | 8 | 10 |
| The hedge war | 8 | 9 |
| Hours cut overnight | 9 | 9 |
| The handwritten will | 9 | 9 |
| £600 and a shrug | 7 | 9 |
| It depends what you signed | 9 | 10 |
| Letter before action | 10 | 10 |
| Refuse the order workaround | 9 | 10 |
Full receipts: GLM 5.2, GPT-5.6 Terra · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra
Questions people ask
Which is better for everyday legal questions: GLM 5.2 or GPT-5.6 Terra?
GPT-5.6 Terra — it scored 93/100 against 86/100 on our 12-task everyday legal questions suite, winning 7 tasks to 0 with 5 tied. Every answer was marked blind by three judges from three rival AI labs.
How was this tested?
Both models answered the identical published everyday legal questions tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.
More everyday legal questions head-to-heads: Claude Opus 4.8 vs GPT-5.6 Terra · Claude Opus 4.8 vs GLM 5.2 · GPT-5.6 Sol vs GPT-5.6 Terra · GLM 5.2 vs GPT-5.6 Sol · GPT-5.3-Codex vs GPT-5.6 Terra · GPT-5.5 vs GPT-5.6 Terra
Full ranking: Best AI for everyday legal questions · model pages: GLM 5.2, GPT-5.6 Terra