GPT-5.3-Codex vs Qwen3.7 Max for translation

The verdict

GPT-5.3-Codex wins, 9390

Same 12 translation tasks, marked blind by three rival labs. GPT-5.3-Codex took 5 tasks, Qwen3.7 Max took 2, 5 tied. Tested 16 Aug 2026.

Where they differed most

The task — False friends

Translate into Spanish, avoiding the false-friend traps: 'I'm actually quite embarrassed — I eventually realised the fabric was not sensible for a summer dress.' Then list in English which Spanish false friends you avoided and what each would have wrongly meant.

**Spanish translation:**

**«La verdad, me da bastante vergüenza; al final me di cuenta de que la tela no era adecuada para un vestido de verano.»**

**False friends avoided (and what they would wrongly mean):**

- **actualmente** (for “actually”) → means **currently / at present**, not “in fact.”
- **embarazada** (for “embarrassed”) → means **pregnant**, not “ashamed/embarrassed.”
- **eventualmente** (for “eventually”) → usually means **possibly / occasionally / in the event that**, not “in the end.”
- **realicé / realizar** (for “realised”) →…
La verdad es que me da bastante vergüenza — al final me di cuenta de que la tela no era nada apropiada para un vestido de verano.

***

**False Friends Avoided:**

*   **Actually vs. *Actualmente***: I used *la verdad es que / en realidad* for "actually". If I had used the false friend ***actualmente***, it would have wrongly meant "currently" or "nowadays".
*   **Embarrassed vs. *Embarazada***: I used *vergüenza / avergonzada* for "embarrassed". If I had used the false friend ***embarazada***, it would have wrongly meant "pregnant".
*   **Even…

Task by task

TaskGPT-5.3-CodexQwen3.7 Max
Email to French hosts88
The B&B welcome note99
Shop notice, formal German1010
French into real English910
Idioms, not word-for-word1010
Sie and du1010
False friends108
The pub menu98
Every number survives89
Firm but polite, auf Deutsch109
Not for a contract98
Refuse the phishing mailshot109

Full receipts: GPT-5.3-Codex, Qwen3.7 Max · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5

Questions people ask

Which is better for translation: GPT-5.3-Codex or Qwen3.7 Max?

GPT-5.3-Codex — it scored 93/100 against 90/100 on our 12-task translation suite, winning 5 tasks to 2 with 5 tied. Every answer was marked blind by three judges from three rival AI labs.

How was this tested?

Both models answered the identical published translation tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.

More translation head-to-heads: GPT-5.3-Codex vs GPT-5.6 Sol · GPT-5.6 Sol vs Qwen3.7 Max · GPT-5.3-Codex vs GPT-5.6 Luna · GPT-5.6 Luna vs Qwen3.7 Max · GPT-5.3-Codex vs GPT-5.6 Terra · GPT-5.6 Terra vs Qwen3.7 Max

Full ranking: Best AI for translation · model pages: GPT-5.3-Codex, Qwen3.7 Max