GLM 5.2 vs GPT-5.6 Luna for everyday legal questions

The verdict

GPT-5.6 Luna wins, 9286

Same 12 everyday legal questions tasks, marked blind by three rival labs. GPT-5.6 Luna took 7 tasks, GLM 5.2 took 1, 4 tied. Tested 16 Aug 2026.

Where they differed most

The task — Translate the clause

A decorator's contract contains: 'The Client shall indemnify and hold harmless the Contractor against all claims, losses, damages and expenses howsoever arising in connection with the Works, save to the extent caused by the Contractor's negligence.' In max 80 words of plain English: what this clause makes the customer promise, the one carve-out it contains, and the question to ask before signing.

You promise to pay for any claims, losses, or damages related to the decorator's work. 

The one carve-out is that the decorator remains responsible for anything caused by their own negligence. 

Before signing, ask: "Does my home insurance cover this broad financial promise?"
The customer promises to reimburse and protect the contractor from any claims, losses, damage or expenses connected with the work, however they arise. The only exception is where the contractor’s own negligence caused them. Before signing, ask: **What specific risks could this make me responsible for, and does my insurance cover them?**

Task by task

TaskGLM 5.2GPT-5.6 Luna
Landlord with a key89.5
The unprotected deposit89
Eight-month-old washing machine99
The scary parking letter99
Translate the clause810
The hedge war89
Hours cut overnight99
The handwritten will99
£600 and a shrug79
It depends what you signed910
Letter before action108
Refuse the order workaround910

Full receipts: GLM 5.2, GPT-5.6 Luna · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra

Questions people ask

Which is better for everyday legal questions: GLM 5.2 or GPT-5.6 Luna?

GPT-5.6 Luna — it scored 92/100 against 86/100 on our 12-task everyday legal questions suite, winning 7 tasks to 1 with 4 tied. Every answer was marked blind by three judges from three rival AI labs.

How was this tested?

Both models answered the identical published everyday legal questions tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.

More everyday legal questions head-to-heads: Claude Opus 4.8 vs GPT-5.6 Luna · Claude Opus 4.8 vs GLM 5.2 · GPT-5.6 Luna vs GPT-5.6 Sol · GLM 5.2 vs GPT-5.6 Sol · GPT-5.6 Luna vs GPT-5.6 Terra · GLM 5.2 vs GPT-5.6 Terra

Full ranking: Best AI for everyday legal questions · model pages: GLM 5.2, GPT-5.6 Luna