GPT-5.6 Sol vs Grok 4.5 for research skills

The verdict

GPT-5.6 Sol wins, 9893

Same 12 research skills tasks, marked blind by three rival labs. GPT-5.6 Sol took 5 tasks, Grok 4.5 took 1, 6 tied. Tested 16 Aug 2026.

Where they differed most

The task — Plan the dig

Someone wants to trace the history of their 1930s semi in Leicester — who built it, who has lived there, what stood on the land before. Lay out the kinds of sources and places in the UK that answer each part (deeds, census-type records, old maps, local archives, newspaper archives), what each one can and cannot tell them, and the sensible order to work in. No URLs — the shape of the search, not links.

A house history is easiest if treated as three overlapping questions:

1. **Who owned or developed the property?**
2. **Who actually lived there?**
3. **How was the site used before the house was built?**

Those questions usually require different records. Also distinguish between the **landowner/developer, architect, building contractor and first purchaser**: “built by” can mean any of them.

## 1. Start by fixing the property’s identity

Before searching names, establish:

- Current full address and postcode
- Title number and parcel boundari…
# Tracing a 1930s Semi in Leicester: Sources, Limits, and Order of Work

Work from the house outward and from the present backward. Confirm what the building *is* and when it appeared before you chase people or earlier land use. Leicester’s interwar suburbs were often built in estates or small speculative rows, so “who built it” may be a firm or local builder rather than a famous architect, and the same sources that date the house often name the developer.

---

## Sensible order of work

1. **Secure the legal and physical identity of the prope…

Task by task

TaskGPT-5.6 SolGrok 4.5
Make it answerable1010
Rank the sources109
Synthesise the disagreement98
Tear down the survey99
Correlation for the neighbour1010
Plan the dig108
What you cannot cite109
Structure the evidence1010
Ask without leading1010
Two sources, head on910
Primary or secondary108
Refuse the invented bibliography1010

Full receipts: GPT-5.6 Sol, Grok 4.5 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5

Questions people ask

Which is better for research skills: GPT-5.6 Sol or Grok 4.5?

GPT-5.6 Sol — it scored 98/100 against 93/100 on our 12-task research skills suite, winning 5 tasks to 1 with 6 tied. Every answer was marked blind by three judges from three rival AI labs.

How was this tested?

Both models answered the identical published research skills tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.

More research skills head-to-heads: DeepSeek V4 Pro vs GPT-5.6 Sol · DeepSeek V4 Pro vs Grok 4.5 · GPT-5.5 vs GPT-5.6 Sol · GPT-5.6 Luna vs GPT-5.6 Sol · GPT-5.3-Codex vs GPT-5.6 Sol · GPT-5.6 Sol vs GPT-5.6 Terra

Full ranking: Best AI for research skills · model pages: GPT-5.6 Sol, Grok 4.5