GPT-5.3-Codex vs GPT-5.5 for travel planning
GPT-5.5 wins, 93–88
Same 12 travel planning tasks, marked blind by three rival labs. GPT-5.5 took 5 tasks, GPT-5.3-Codex took 1, 6 tied. Tested 16 Aug 2026.
Where they differed most
A wheelchair user (self-propelled chair, can manage one or two steps with help) is planning a day in York: the Minster, the city walls, a riverboat cruise, afternoon tea, the Shambles. For each item, state what specifically must be checked before booking — the exact question to ask, not 'check accessibility' — then reorder the day so the item most at risk of disappointing comes first while energy is highest, naming which one that is.
Great plan. Here are the **exact questions** to ask for each stop. ## 1) York Minster — questions to ask 1. “Which entrance is step-free for a self-propelled wheelchair, and what postcode/what3words should I use for that entrance?” 2. “On the standard visitor route, are there any unavoidable steps, and if so exactly where?” 3. “Can I reach the main floor, Undercroft museum, shop, and café without stairs?” 4. “Is there an accessible toilet on-site, and does it require a RADAR key?” 5. “Do you offer a free companion/carer ticket, and how do I bo…
## Most at risk of disappointing: **York City Walls** For a self-propelled wheelchair user, the **city walls** are the item most likely to disappoint because access often involves flights of steps, narrow/uneven wall-walk sections, and limited step-free options. Put this first while energy and patience are highest — and be ready to switch to a ground-level “walls and bars” route if the wall-walk itself is not feasible. --- ## What to ask before booking / committing ### 1) York Minster Ask: **“On the date and time we plan to visit, is there…
Task by task
| Task | GPT-5.3-Codex | GPT-5.5 |
|---|---|---|
| The impossible family holiday | 8 | 9 |
| Three days, one toddler | 9 | 9 |
| Hand luggage, February, Krakow | 9 | 9 |
| Will we make it? | 8 | 9 |
| Total the real cost | 8 | 9 |
| The ferry is cancelled | 9 | 9 |
| Check, don't assume | 8 | 10 |
| Declare it or regret it | 10 | 9 |
| Next September's weather | 8 | 10 |
| Brief the dog-sitter | 10 | 10 |
| Tokyo by Thursday, sharp by Friday | 9 | 9 |
| Refuse the customs dodge | 10 | 10 |
Full receipts: GPT-5.3-Codex, GPT-5.5 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5
Questions people ask
Which is better for travel planning: GPT-5.3-Codex or GPT-5.5?
GPT-5.5 — it scored 93/100 against 88/100 on our 12-task travel planning suite, winning 5 tasks to 1 with 6 tied. Every answer was marked blind by three judges from three rival AI labs.
How was this tested?
Both models answered the identical published travel planning tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.
More travel planning head-to-heads: GPT-5.5 vs GPT-5.6 Sol · GPT-5.3-Codex vs GPT-5.6 Sol · GPT-5.5 vs GPT-5.6 Luna · GPT-5.3-Codex vs GPT-5.6 Luna · Claude Fable 5 vs GPT-5.5 · GPT-5.5 vs GPT-5.6 Terra
Full ranking: Best AI for travel planning · model pages: GPT-5.3-Codex, GPT-5.5