GPT-5.3-Codex vs Kimi K3 for translation

The verdict

GPT-5.3-Codex wins, 9387

Same 12 translation tasks, marked blind by three rival labs. GPT-5.3-Codex took 8 tasks, Kimi K3 took 2, 2 tied. Tested 16 Aug 2026.

Where they differed most

The task — Shop notice, formal German

Translate into German, formal Sie register, reading like a real German shop notice — warm but proper: 'Dear customers, from Monday 2 March we will open at 7.30am instead of 7am. The bread delivery for pre-orders is unaffected. Thank you for your understanding — we know how important the early loaves are for some of you. The Hartley family.'

**Liebe Kundinnen und Kunden,**

ab **Montag, dem 2. März**, öffnen wir statt um **7:00 Uhr** erst um **7:30 Uhr**.  
Die Brotauslieferung für Vorbestellungen bleibt davon unberührt.

Vielen Dank für Ihr Verständnis – wir wissen, wie wichtig einigen von Ihnen die frühen Brote sind.

**Ihre Familie Hartley**
**Liebe Kundinnen und Kunden,**

ab Montag, dem 2. März, öffnen wir um 7.30 Uhr statt wie bisher um 7.00 Uhr.

Die Auslieferung vorbestellter Brote bleibt davon unberührt.

Vielen Dank für Ihr Verständnis – wir wissen, wie wichtig für manche von Ihnen das frische Brot am frühen Morgen ist.

**Ihre Familie Hartley**

A few notes on the choices:
- "Liebe Kundinnen und Kunden" is the standard warm-but-proper opening on German shop notices (Aushänge); if you'd like it a touch more formal, "Sehr geehrte Kundinnen und Kunden" works too.
- "7.30 Uhr"…

Task by task

TaskGPT-5.3-CodexKimi K3
Email to French hosts88
The B&B welcome note98
Shop notice, formal German108
French into real English98
Idioms, not word-for-word109
Sie and du109
False friends1010
The pub menu98
Every number survives810
Firm but polite, auf Deutsch108
Not for a contract910
Refuse the phishing mailshot108

Full receipts: GPT-5.3-Codex, Kimi K3 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5

Questions people ask

Which is better for translation: GPT-5.3-Codex or Kimi K3?

GPT-5.3-Codex — it scored 93/100 against 87/100 on our 12-task translation suite, winning 8 tasks to 2 with 2 tied. Every answer was marked blind by three judges from three rival AI labs.

How was this tested?

Both models answered the identical published translation tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.

More translation head-to-heads: GPT-5.3-Codex vs GPT-5.6 Sol · GPT-5.6 Sol vs Kimi K3 · GPT-5.3-Codex vs GPT-5.6 Luna · GPT-5.6 Luna vs Kimi K3 · GPT-5.3-Codex vs GPT-5.6 Terra · GPT-5.6 Terra vs Kimi K3

Full ranking: Best AI for translation · model pages: GPT-5.3-Codex, Kimi K3