Claude Opus 4.8 vs GPT-5.3-Codex for translation
GPT-5.3-Codex wins, 93–90
Same 12 translation tasks, marked blind by three rival labs. GPT-5.3-Codex took 3 tasks, Claude Opus 4.8 took 2, 7 tied. Tested 16 Aug 2026.
Where they differed most
Translate this pub menu into French for a chalkboard, adding a short French gloss in brackets wherever a dish has no French equivalent — a French visitor should end up knowing what they would actually be eating: 'Toad in the hole with onion gravy. Ploughman's lunch. Sticky toffee pudding with custard. Scotch egg. Mushy peas.'
Here's a chalkboard-ready translation with explanatory glosses for the dishes that have no French equivalent: --- **Toad in the Hole, sauce à l'oignon** *(saucisses rôties dans une pâte à crêpe salée cuite au four, nappées d'une sauce brune à l'oignon)* **Ploughman's Lunch** *(assiette froide du laboureur : fromage, pain, cornichons, pickles et salade)* **Sticky Toffee Pudding, crème anglaise** *(moelleux aux dattes nappé de caramel, servi avec crème anglaise)* **Scotch Egg** *(œuf dur enrobé de chair à saucisse, pané et frit)* **Mushy Pe…
- **Toad in the hole, sauce aux oignons** *[saucisses cuites au four dans une pâte type Yorkshire pudding]* - **Déjeuner du laboureur (Ploughman’s lunch)** *[assiette froide avec pain, fromage, pickles/chutney, parfois charcuterie]* - **Sticky toffee pudding, crème anglaise** *[gâteau moelleux (souvent aux dattes) nappé de sauce toffee/caramel]* - **Œuf à l’écossaise** *[œuf dur enrobé de chair à saucisse, pané puis frit]* - **Purée de petits pois**
Task by task
| Task | Claude Opus 4.8 | GPT-5.3-Codex |
|---|---|---|
| Email to French hosts | 8 | 8 |
| The B&B welcome note | 9 | 9 |
| Shop notice, formal German | 9 | 10 |
| French into real English | 9 | 9 |
| Idioms, not word-for-word | 10 | 10 |
| Sie and du | 10 | 10 |
| False friends | 10 | 10 |
| The pub menu | 5 | 9 |
| Every number survives | 10 | 8 |
| Firm but polite, auf Deutsch | 8 | 10 |
| Not for a contract | 10 | 9 |
| Refuse the phishing mailshot | 10 | 10 |
Full receipts: Claude Opus 4.8, GPT-5.3-Codex · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5
Questions people ask
Which is better for translation: Claude Opus 4.8 or GPT-5.3-Codex?
GPT-5.3-Codex — it scored 93/100 against 90/100 on our 12-task translation suite, winning 3 tasks to 2 with 7 tied. Every answer was marked blind by three judges from three rival AI labs.
How was this tested?
Both models answered the identical published translation tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.
More translation head-to-heads: GPT-5.3-Codex vs GPT-5.6 Sol · Claude Opus 4.8 vs GPT-5.6 Sol · GPT-5.3-Codex vs GPT-5.6 Luna · Claude Opus 4.8 vs GPT-5.6 Luna · GPT-5.3-Codex vs GPT-5.6 Terra · Claude Opus 4.8 vs GPT-5.6 Terra
Full ranking: Best AI for translation · model pages: Claude Opus 4.8, GPT-5.3-Codex