GPT-5.3-Codex vs GPT-5.5 for travel planning

The verdict

GPT-5.5 wins, 9388

Same 12 travel planning tasks, marked blind by three rival labs. GPT-5.5 took 5 tasks, GPT-5.3-Codex took 1, 6 tied. Tested 16 Aug 2026.

Where they differed most

The task — Check, don't assume

A wheelchair user (self-propelled chair, can manage one or two steps with help) is planning a day in York: the Minster, the city walls, a riverboat cruise, afternoon tea, the Shambles. For each item, state what specifically must be checked before booking — the exact question to ask, not 'check accessibility' — then reorder the day so the item most at risk of disappointing comes first while energy is highest, naming which one that is.

Great plan. Here are the **exact questions** to ask for each stop.

## 1) York Minster — questions to ask
1. “Which entrance is step-free for a self-propelled wheelchair, and what postcode/what3words should I use for that entrance?”
2. “On the standard visitor route, are there any unavoidable steps, and if so exactly where?”
3. “Can I reach the main floor, Undercroft museum, shop, and café without stairs?”
4. “Is there an accessible toilet on-site, and does it require a RADAR key?”
5. “Do you offer a free companion/carer ticket, and how do I bo…
GPT-5.510/10
## Most at risk of disappointing: **York City Walls**
For a self-propelled wheelchair user, the **city walls** are the item most likely to disappoint because access often involves flights of steps, narrow/uneven wall-walk sections, and limited step-free options. Put this first while energy and patience are highest — and be ready to switch to a ground-level “walls and bars” route if the wall-walk itself is not feasible.

---

## What to ask before booking / committing

### 1) York Minster
Ask:  
**“On the date and time we plan to visit, is there…

Task by task

TaskGPT-5.3-CodexGPT-5.5
The impossible family holiday89
Three days, one toddler99
Hand luggage, February, Krakow99
Will we make it?89
Total the real cost89
The ferry is cancelled99
Check, don't assume810
Declare it or regret it109
Next September's weather810
Brief the dog-sitter1010
Tokyo by Thursday, sharp by Friday99
Refuse the customs dodge1010

Full receipts: GPT-5.3-Codex, GPT-5.5 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5

Questions people ask

Which is better for travel planning: GPT-5.3-Codex or GPT-5.5?

GPT-5.5 — it scored 93/100 against 88/100 on our 12-task travel planning suite, winning 5 tasks to 1 with 6 tied. Every answer was marked blind by three judges from three rival AI labs.

How was this tested?

Both models answered the identical published travel planning tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.

More travel planning head-to-heads: GPT-5.5 vs GPT-5.6 Sol · GPT-5.3-Codex vs GPT-5.6 Sol · GPT-5.5 vs GPT-5.6 Luna · GPT-5.3-Codex vs GPT-5.6 Luna · Claude Fable 5 vs GPT-5.5 · GPT-5.5 vs GPT-5.6 Terra

Full ranking: Best AI for travel planning · model pages: GPT-5.3-Codex, GPT-5.5