Best AI for travel planning

The verdict

GPT-5.6 Sol

Scored 94/100 on our 12-task travel-planning suitea single point ahead of GPT-5.6 Luna (93) — effectively level.

What to actually do

GPT-5.6 Sol is the engine inside ChatGPT. Go to chatgpt.comthe free tier is fine to start. Paid plans start at £7/month (go plan, vendor’s own price). The free tier is ChatGPT’s, not a promise about this exact model — we haven’t verified which plan carries it. Not fussed about the last point or two? Any of the top 4 here will serve you well.

Real trips are constraint problems: five people, £1,600, a gran who can't walk far and a 9-year-old obsessed with castles. The suite tests planning judgement on material given in the task — itinerary logic with a toddler, whether a 14:20 landing makes a 16:00 coach, totalling a France trip honestly, and what to do the night the ferry is cancelled. One task asks how to hide prescription codeine from Dubai customs — the right answer refuses and points at the legitimate route.

updated 16 Aug 2026 · tested by Robert Prime · re-ranks automatically when a new run lands

#ModelOur score
1GPT-5.6 Sollatest94/100
2GPT-5.6 Lunalatest93/100
3GPT-5.593/100
4Claude Fable 593/100
5GPT-5.6 Terra91/100
6Claude Sonnet 588/100
7GPT-5.3-Codex88/100
8Claude Opus 4.888/100
9Claude Opus 4.686/100
10Grok 4.583/100
11DeepSeek V4 Pro82/100
12Kimi K382/100
13Qwen3.7 Max81/100
14Gemini 3.5 Flash80/100
15Gemini 3.1 Pro Preview79/100
16Gemini 3.1 Flash Lite76/100
17GLM 5.276/100
18DeepSeek V4 Flash68/100
19Mistral Medium 3.568/100

“API cost” is what software developers pay to build on a model — ignore it if you just use the website. Each model answers each task once. Models level on score are ranked by a fixed tie-break — fewest machine-checked rule breaches, then lowest measured cost per run — so the order is deterministic and checkable, never arbitrary. Judge panels never include the contestant’s own lab, so panels differ slightly per model — small cross-model gaps can reflect panel severity, not quality.

1.

GPT-5.6 Sol

94/100our pick — won on the tie-break

Made by OpenAI — their newest model. You use it inside ChatGPT — nothing to install.

Strongest showing: Three days, one toddler” — scored 9/10 by the panel.

Accurate plan using only listed sites; one anchor, nap rule, rain plans daily; drops Camera Obscura correctly for stairs. Clear, useful, minor polish only.x-ai/grok-4.5, judging blind · full receipts ↓

12 tasks · 16 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $5 in / $30 out per 1M tokens · full model page →

2.

GPT-5.6 Luna

93/100

Made by OpenAI — their newest model. You use it inside ChatGPT — nothing to install.

Strongest showing: Three days, one toddler” — scored 9/10 by the panel.

Excellent toddler-friendly plan using only listed attractions, one anchor daily, post-nap timing, rain alternatives, and clear drop rationale for Camera Obscura.x-ai/grok-4.5, judging blind · full receipts ↓

12 tasks · 16 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $0.1 in / $0.6 out per 1M tokens · full model page →

3.

GPT-5.5

93/100

Made by OpenAI. You use it inside ChatGPT — nothing to install.

Strongest showing: Check, don't assume” — scored 10/10 by the panel.

The response perfectly follows all instructions, providing highly specific and practical questions for each attraction. It correctly identifies the City Walls as the most challenging, reorders the itinerary logically, and offers excellent, google/gemini-3.1-pro-preview, judging blind · full receipts ↓

12 tasks · 16 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $5 in / $30 out per 1M tokens · full model page →

See it for yourself

One task from the suite, and what the leader, the middle of the table and the bottom actually wrote. This is the task they disagreed on most — scores ranged from 5 to 9 out of 10.

The task — The impossible family holiday

Plan the SHAPE of a holiday from these constraints, all real: family of five — two parents, kids 9 and 6, and gran who cannot walk more than ten minutes; £1,600 all-in for a week in August; flying from Manchester only; dad burns easily and hates 'just sitting on a beach'; the 9-year-old is obsessed with castles. Offer 2 distinct holiday shapes (types of place and trip, not named resorts), what each one trades away, and the single constraint you would challenge first because it costs the most.

At £1,600 for five, the trip has to average about **£46 per person per day, including flights**. Both shapes therefore assume early booking, self-catering, minimal luggage, public transport and few paid attractions.

## 1. Compact, castle-rich Central/Eastern European city

**Shape:**  
Seven nights…

Judge: Meets all task elements accurately with two viable shapes, clear trade-offs, and smart constraint challenge; highly useful and polished with only minor length.

**Two holiday shapes**

**Shape 1: Compact historic-city base (short-haul flight + one accessible apartment)**  
Fly from Manchester to a single walkable/ fortifiable northern or central-European city with major castles or fortified old towns. Stay the full week in one ground-floor or lift-served se…

Judge: Well-structured, addresses all constraints, offers two distinct shapes with tradeoffs, and justifies budget as top constraint. Slightly verbose but clear and useful.

Here are two distinct holiday shapes that fit your constraints, each with trade-offs, plus the constraint I’d challenge first:

---

### **1. City Break with Castle Focus (e.g., Edinburgh or Kraków)**
**Shape:**
- Fly from Manchester to a historic city with castles (e.g., Edinburgh Castle, Wawel Cas…

Judge: Task asked for holiday 'shapes' not named resorts, but response names specific cities/castles, violating instruction. Flying-from-Manchester challenge is odd since task states it as fixed 'real' constraint. Useful structure but factual/logi

Trimmed to fit. Every full answer, every judge’s score: GPT-5.6 Sol, Grok 4.5, Mistral Medium 3.5

How this ranking is made

Every model answers the identical published task suite. Three AI judges from three different labs mark each answer blind against a fixed 0–10 rubric — none is ever from the maker’s own lab, and an answer that breaks a task’s explicit rules is capped in code, not by the judge’s goodwill. The score is the average, out of 100.

What this is not: an opinion piece, an affiliate list, or a benchmark we cannot show you. Every score links its raw outputs and judge verdicts. The full protocol · receipts: GPT-5.6 Sol, GPT-5.6 Luna, GPT-5.5, Claude Fable 5, GPT-5.6 Terra, Claude Sonnet 5, GPT-5.3-Codex, Claude Opus 4.8, Claude Opus 4.6, Grok 4.5, DeepSeek V4 Pro, Kimi K3, Qwen3.7 Max, Gemini 3.5 Flash, Gemini 3.1 Pro Preview, Gemini 3.1 Flash Lite, GLM 5.2, DeepSeek V4 Flash, Mistral Medium 3.5

Questions people ask

What is the best AI for travel planning in 2026?

GPT-5.6 Sol leads our tested ranking with 94/100 on our 12-task travel-planning suite (12 tasks), in a dead heat with GPT-5.6 Luna (93). Every answer was marked blind by three AI judges from three different labs, and the full outputs are downloadable.

How is this ranking made?

Each model answers the identical published task suite; three judges from different labs score every answer 0–10 against a fixed rubric without knowing which produced it; answers that break a task's explicit rules are capped automatically. The score is the average, out of 100. No vendor pays for placement.

How often does this page update?

It re-ranks itself whenever a new test run lands, and prices re-verify daily against vendor pages. The current ranking was last computed on 16 Aug 2026.

Head-to-heads in travel planning

Show all 20 tested pairs ▾

All comparisons →

More rankings ▾

Best AI for writing · Best AI chatbot for everyday use · Best AI for coding · Best free AI model · Best AI for spreadsheets and Excel · Best AI essay writer · Best AI for summarising documents · Best AI for extracting data from text · Best AI for writing emails · Best AI for everyday maths and percentages · Best AI for customer service replies · Best AI for revision and study notes · Best AI for vibe coding · Best AI for making flashcards · Best AI for social media posts · Best AI for job applications and cover letters · Best AI for presentations · Best AI for writing your CV · Best AI for research skills · Best AI for creative writing · Best AI for translation · Best AI for emotional support · Best AI for everyday legal questions · Best AI for health questions · Best AI for writing a book · Best AI for humanising AI text · Best-value AI model API · every model we track · every tool