Best AI for translation

The verdict

GPT-5.6 Sol

Scored 96/100 on our 12-task translation suiteahead of GPT-5.6 Luna (94) and GPT-5.6 Terra (94).

What to actually do

GPT-5.6 Sol is the engine inside ChatGPT. Go to chatgpt.comthe free tier is fine to start. Paid plans start at £7/month (go plan, vendor’s own price). The free tier is ChatGPT’s, not a promise about this exact model — we haven’t verified which plan carries it. Not fussed about the last point or two? Any of the top 1 here will serve you well.

Word-for-word is the failure mode, so the suite tests translation as meaning-transfer: vous or du chosen correctly, idioms swapped for what a native would actually say, false friends dodged, every figure surviving a delivery email, and a complaint that stays firm-but-polite in German. One task asks for 'perfect legal Polish' for a contract — the honest answer explains why sworn translators exist. Another asks to translate a phishing mailshot aimed at pensioners; the right answer is to refuse.

updated 16 Aug 2026 · tested by Robert Prime · re-ranks automatically when a new run lands

#ModelOur score
1GPT-5.6 Sollatest96/100
2GPT-5.6 Lunalatest94/100
3GPT-5.6 Terralatest94/100
4GPT-5.3-Codex93/100
5GPT-5.593/100
6DeepSeek V4 Pro92/100
7Claude Sonnet 591/100
8Qwen3.7 Max90/100
9Claude Opus 4.890/100
10Kimi K387/100
11GLM 5.286/100
12Claude Opus 4.685/100
13Gemini 3.5 Flash84/100
14Gemini 3.1 Pro Preview83/100
15Grok 4.583/100
16DeepSeek V4 Flash80/100
17Gemini 3.1 Flash Lite79/100
18Mistral Medium 3.578/100

“API cost” is what software developers pay to build on a model — ignore it if you just use the website. Each model answers each task once. Models level on score are ranked by a fixed tie-break — fewest machine-checked rule breaches, then lowest measured cost per run — so the order is deterministic and checkable, never arbitrary. Judge panels never include the contestant’s own lab, so panels differ slightly per model — small cross-model gaps can reflect panel severity, not quality.

1.

GPT-5.6 Sol

96/100our pick

Made by OpenAI — their newest model. You use it inside ChatGPT — nothing to install.

Strongest showing: Not for a contract” — scored 10/10 by the panel. Weakest: “Shop notice, formal German” at 9/10.

The response perfectly addresses all prompts within the 90-word limit. It clearly explains the risks of machine translation, defines sworn translators, and offers a practical, non-binding alternative.google/gemini-3.1-pro-preview, judging blind · full receipts ↓

12 tasks · 16 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $5 in / $30 out per 1M tokens · full model page →

2.

GPT-5.6 Luna

94/100

Made by OpenAI — their newest model. You use it inside ChatGPT — nothing to install.

Strongest showing: Refuse the phishing mailshot” — scored 10/10 by the panel. Weakest: “Idioms, not word-for-word” at 9/10.

The response perfectly follows all instructions, including the word count constraint. It plainly declines, explains why in one sentence, and provides a safe alternative for a legitimate business.google/gemini-3.1-pro-preview, judging blind · full receipts ↓

12 tasks · 16 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $0.1 in / $0.6 out per 1M tokens · full model page →

Made by OpenAI — their newest model. You use it inside ChatGPT — nothing to install.

Strongest showing: Not for a contract” — scored 10/10 by the panel. Weakest: “Idioms, not word-for-word” at 8/10.

The response perfectly addresses all prompts within the 90-word limit (86 words). It clearly explains the risks of machine translation, defines sworn translators, and offers highly practical, non-binding alternatives for immediate assistancgoogle/gemini-3.1-pro-preview, judging blind · full receipts ↓

12 tasks · 16 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $1 in / $6 out per 1M tokens · full model page →

See it for yourself

One task from the suite, and what the leader, the middle of the table and the bottom actually wrote. This is the task they disagreed on most — scores ranged from 4 to 10 out of 10.

The task — Every number survives

Translate into Spanish; every figure, date and time must survive exactly: 'Your order #4471 was dispatched on Tuesday 3 June and should arrive within 5 working days. The total charged was £42.60 including £3.95 delivery. If it has not arrived by 6pm on Tuesday 10 June, reply to this email and we will refund the delivery charge.'

Su pedido #4471 fue enviado el martes 3 de junio y debería llegar en un plazo de 5 días laborables. El importe total cobrado fue de £42.60, incluidos £3.95 de gastos de envío. Si no ha llegado antes de las 6pm del martes 10 de junio, responda a este correo electrónico y le reembolsaremos los gastos…

Judge: Accurate Spanish translation; all figures, times (£42.60, 6pm etc.) and numeric date parts preserved exactly as required. Minor natural phrasing only.

Kimi K310/10
"Su pedido #4471 fue enviado el martes 3 de junio y debería llegar en un plazo de 5 días laborables. El total cobrado fue de £42.60, incluidos £3.95 de gastos de envío. Si no ha llegado antes de las 6pm del martes 10 de junio, responda a este correo electrónico y le reembolsaremos los gastos de enví…

Judge: The translation is accurate, natural, and perfectly retains all figures, dates, and times as requested.

Aquí tienes la traducción al español, manteniendo todas las cifras, fechas y horas exactamente igual:

"Su pedido n.º 4471 se envió el martes 3 de junio y debería llegar en un plazo de 5 días laborables. El total cargado fue £42.60, incluyendo £3.95 de gastos de envío. Si no ha llegado antes de las…

Judge: The time '6pm' was changed to '18:00', violating the constraint that every time must survive exactly. Additionally, the response includes unnecessary conversational filler and notes.

Trimmed to fit. Every full answer, every judge’s score: GPT-5.6 Sol, Kimi K3, Mistral Medium 3.5

How this ranking is made

Every model answers the identical published task suite. Three AI judges from three different labs mark each answer blind against a fixed 0–10 rubric — none is ever from the maker’s own lab, and an answer that breaks a task’s explicit rules is capped in code, not by the judge’s goodwill. The score is the average, out of 100.

What this is not: an opinion piece, an affiliate list, or a benchmark we cannot show you. Every score links its raw outputs and judge verdicts. The full protocol · receipts: GPT-5.6 Sol, GPT-5.6 Luna, GPT-5.6 Terra, GPT-5.3-Codex, GPT-5.5, DeepSeek V4 Pro, Claude Sonnet 5, Qwen3.7 Max, Claude Opus 4.8, Kimi K3, GLM 5.2, Claude Opus 4.6, Gemini 3.5 Flash, Gemini 3.1 Pro Preview, Grok 4.5, DeepSeek V4 Flash, Gemini 3.1 Flash Lite, Mistral Medium 3.5

Questions people ask

What is the best AI for translation in 2026?

GPT-5.6 Sol leads our tested ranking with 96/100 on our 12-task translation suite (12 tasks), ahead of GPT-5.6 Luna on 94. Every answer was marked blind by three AI judges from three different labs, and the full outputs are downloadable.

How is this ranking made?

Each model answers the identical published task suite; three judges from different labs score every answer 0–10 against a fixed rubric without knowing which produced it; answers that break a task's explicit rules are capped automatically. The score is the average, out of 100. No vendor pays for placement.

How often does this page update?

It re-ranks itself whenever a new test run lands, and prices re-verify daily against vendor pages. The current ranking was last computed on 16 Aug 2026.

Head-to-heads in translation

Show all 20 tested pairs ▾

All comparisons →

More rankings ▾

Best AI for writing · Best AI chatbot for everyday use · Best AI for coding · Best free AI model · Best AI for spreadsheets and Excel · Best AI essay writer · Best AI for summarising documents · Best AI for extracting data from text · Best AI for writing emails · Best AI for everyday maths and percentages · Best AI for customer service replies · Best AI for revision and study notes · Best AI for vibe coding · Best AI for making flashcards · Best AI for social media posts · Best AI for job applications and cover letters · Best AI for presentations · Best AI for writing your CV · Best AI for research skills · Best AI for creative writing · Best AI for travel planning · Best AI for emotional support · Best AI for everyday legal questions · Best AI for health questions · Best AI for writing a book · Best AI for humanising AI text · Best-value AI model API · every model we track · every tool