GPT-5.6 Luna vs Kimi K3 for translation

The verdict

GPT-5.6 Luna wins, 9487

Same 12 translation tasks, marked blind by three rival labs. GPT-5.6 Luna took 7 tasks, Kimi K3 took 2, 3 tied. Tested 16 Aug 2026.

Where they differed most

The task — Shop notice, formal German

Translate into German, formal Sie register, reading like a real German shop notice — warm but proper: 'Dear customers, from Monday 2 March we will open at 7.30am instead of 7am. The bread delivery for pre-orders is unaffected. Thank you for your understanding — we know how important the early loaves are for some of you. The Hartley family.'

**Liebe Kundinnen und Kunden,**

ab Montag, dem 2. März, öffnen wir erst um 7:30 Uhr statt um 7:00 Uhr.

Die Lieferung der vorbestellten Brote bleibt davon unberührt.

Vielen Dank für Ihr Verständnis – wir wissen, wie wichtig die Brote am frühen Morgen für einige von Ihnen sind.

**Ihre Familie Hartley**
**Liebe Kundinnen und Kunden,**

ab Montag, dem 2. März, öffnen wir um 7.30 Uhr statt wie bisher um 7.00 Uhr.

Die Auslieferung vorbestellter Brote bleibt davon unberührt.

Vielen Dank für Ihr Verständnis – wir wissen, wie wichtig für manche von Ihnen das frische Brot am frühen Morgen ist.

**Ihre Familie Hartley**

A few notes on the choices:
- "Liebe Kundinnen und Kunden" is the standard warm-but-proper opening on German shop notices (Aushänge); if you'd like it a touch more formal, "Sehr geehrte Kundinnen und Kunden" works too.
- "7.30 Uhr"…

Task by task

TaskGPT-5.6 LunaKimi K3
Email to French hosts98
The B&B welcome note98
Shop notice, formal German108
French into real English108
Idioms, not word-for-word99
Sie and du109
False friends1010
The pub menu88
Every number survives910
Firm but polite, auf Deutsch108
Not for a contract910
Refuse the phishing mailshot108

Full receipts: GPT-5.6 Luna, Kimi K3 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5

Questions people ask

Which is better for translation: GPT-5.6 Luna or Kimi K3?

GPT-5.6 Luna — it scored 94/100 against 87/100 on our 12-task translation suite, winning 7 tasks to 2 with 3 tied. Every answer was marked blind by three judges from three rival AI labs.

How was this tested?

Both models answered the identical published translation tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.

More translation head-to-heads: GPT-5.6 Luna vs GPT-5.6 Sol · GPT-5.6 Sol vs Kimi K3 · GPT-5.6 Luna vs GPT-5.6 Terra · GPT-5.3-Codex vs GPT-5.6 Luna · GPT-5.5 vs GPT-5.6 Luna · DeepSeek V4 Pro vs GPT-5.6 Luna

Full ranking: Best AI for translation · model pages: GPT-5.6 Luna, Kimi K3