GPT-5.5 vs Kimi K3 for research skills
GPT-5.5 wins, 97–90
Same 12 research skills tasks, marked blind by three rival labs. GPT-5.5 took 5 tasks, Kimi K3 took 1, 6 tied. Tested 16 Aug 2026.
Where they differed most
Someone wants to trace the history of their 1930s semi in Leicester — who built it, who has lived there, what stood on the land before. Lay out the kinds of sources and places in the UK that answer each part (deeds, census-type records, old maps, local archives, newspaper archives), what each one can and cannot tell them, and the sensible order to work in. No URLs — the shape of the search, not links.
Below is a practical “shape of the search” for a 1930s semi-detached house in Leicester: who built it, who lived there, and what was on the land before. The key is to work from the known present address backwards, while allowing for street renumbering, ward/parish boundary changes, and estate-name changes. ## 1. Start by fixing the identity of the house Before chasing records, establish: - Current full address and postcode. - Title number, if registered. - Whether the street has ever been renamed or renumbered. - Whether the house was origin…
# Tracing a 1930s semi in Leicester Good news first: a 1930s house sits in a sweet spot. It's recent enough that paper survives in quantity (deeds, building control plans, the 1939 Register), but old enough that the land underneath was almost certainly fields, so there's a "before" worth finding. One early fork to resolve: is it a **private spec-built semi** (Knighton, Evington, Rushey Mead, Western Park and similar) or an **ex-council house** (Braunstone, New Parks, Saffron Lane estates)? The deeds will tell you immediately — if the first con…
Task by task
| Task | GPT-5.5 | Kimi K3 |
|---|---|---|
| Make it answerable | 10 | 10 |
| Rank the sources | 9 | 9 |
| Synthesise the disagreement | 9 | 9 |
| Tear down the survey | 10 | 9 |
| Correlation for the neighbour | 9 | 10 |
| Plan the dig | 9 | 5 |
| What you cannot cite | 10 | 10 |
| Structure the evidence | 10 | 10 |
| Ask without leading | 10 | 9 |
| Two sources, head on | 10 | 8 |
| Primary or secondary | 10 | 9 |
| Refuse the invented bibliography | 10 | 10 |
Full receipts: GPT-5.5, Kimi K3 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5
Questions people ask
Which is better for research skills: GPT-5.5 or Kimi K3?
GPT-5.5 — it scored 97/100 against 90/100 on our 12-task research skills suite, winning 5 tasks to 1 with 6 tied. Every answer was marked blind by three judges from three rival AI labs.
How was this tested?
Both models answered the identical published research skills tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.
More research skills head-to-heads: DeepSeek V4 Pro vs GPT-5.5 · DeepSeek V4 Pro vs Kimi K3 · GPT-5.5 vs GPT-5.6 Sol · GPT-5.6 Sol vs Kimi K3 · GPT-5.5 vs GPT-5.6 Luna · GPT-5.3-Codex vs GPT-5.5
Full ranking: Best AI for research skills · model pages: GPT-5.5, Kimi K3