GPT-5.5 vs GPT-5.6 Terra for everyday legal questions

The verdict

Dead heat: 93 each

Same 12 everyday legal questions tasks, marked blind by three rival labs. Tested 16 Aug 2026.

Where they differed most

The task — Translate the clause

A decorator's contract contains: 'The Client shall indemnify and hold harmless the Contractor against all claims, losses, damages and expenses howsoever arising in connection with the Works, save to the extent caused by the Contractor's negligence.' In max 80 words of plain English: what this clause makes the customer promise, the one carve-out it contains, and the question to ask before signing.

You promise to cover the decorator for any claims, losses, damage, or costs connected with the work, even if they come from third parties. The only exception is where the loss was caused by the decorator’s own negligence. Before signing, ask: “Am I comfortable taking on such broad liability, and should it be limited to losses I cause or that my insurance covers?”
The customer promises to cover the decorator’s costs and liabilities for any claims, losses, damage or expenses connected with the work. The exception is where those costs were caused by the decorator’s own negligence. Before signing, ask: “Why am I taking responsibility for all work-related claims, including risks I may not control, and can this be limited to losses caused by me?”

Task by task

TaskGPT-5.5GPT-5.6 Terra
Landlord with a key88
The unprotected deposit99
Eight-month-old washing machine1010
The scary parking letter99
Translate the clause910
The hedge war99
Hours cut overnight99
The handwritten will99
£600 and a shrug99
It depends what you signed1010
Letter before action1010
Refuse the order workaround1010

Full receipts: GPT-5.5, GPT-5.6 Terra · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5

Questions people ask

Which is better for everyday legal questions: GPT-5.5 or GPT-5.6 Terra?

They tied at 93/100 on our everyday legal questions suite — for this job, pick on price or the product you already use.

How was this tested?

Both models answered the identical published everyday legal questions tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.

More everyday legal questions head-to-heads: Claude Opus 4.8 vs GPT-5.6 Terra · Claude Opus 4.8 vs GPT-5.5 · GPT-5.6 Sol vs GPT-5.6 Terra · GPT-5.5 vs GPT-5.6 Sol · GPT-5.3-Codex vs GPT-5.6 Terra · Claude Fable 5 vs GPT-5.6 Terra

Full ranking: Best AI for everyday legal questions · model pages: GPT-5.5, GPT-5.6 Terra