GPT-5.3-Codex vs Kimi K3 for workflow automation
GPT-5.3-Codex wins, 90–85
Same 12 workflow automation tasks, marked blind by three rival labs. GPT-5.3-Codex took 6 tasks, Kimi K3 took 1, 5 tied. Tested 27 Aug 2026.
Where they differed most
The nightly job that imports supplier price files has failed. What is true: it last succeeded 26 hours ago; the shop is still selling at yesterday's prices; three suppliers changed prices overnight; re-running the import by hand takes about 20 minutes; running it twice would double-apply a percentage uplift and corrupt the price list. Write ONLY the alert message that goes to the on-call person at 3am — max 60 words, nothing else, no heading, no explanation of your choices. It must say what has broken, what it is costing while it stays broken, the first action to take, and the one thing they must not do.
Supplier price import failed; last successful run was 26h ago. We’re still selling at yesterday’s prices while three suppliers changed prices overnight, so current prices/margins are wrong. First action: manually run the import once now (about 20 minutes) and confirm success. Do NOT run it twice—second run double-applies uplift and corrupts the price list.
Supplier price import failed overnight (last success 26h ago). Shop is still selling at yesterday's prices; three suppliers changed theirs, so margin is wrong on those lines until fixed. Re-run the import manually once — about 20 minutes. Do NOT run it twice: a second run double-applies the uplift and corrupts the price list.
Task by task
| Task | GPT-5.3-Codex | Kimi K3 |
|---|---|---|
| The MOT reminder nobody sends | 9 | 8 |
| The spec, as JSON | 9 | 8 |
| Where the human stays | 9 | 9 |
| The unhappy paths are the job | 8 | 8 |
| Do not automate this | 8 | 9 |
| It fired twice | 9 | 9 |
| The spreadsheet that runs the business | 9 | 8 |
| The alert that wakes someone up | 10 | 8 |
| How long it really takes | 9 | 9 |
| No API, no chance? | 9 | 8 |
| Rules, not guesses | 9 | 9 |
| Refuse the scraper | 10 | 9 |
Full receipts: GPT-5.3-Codex, Kimi K3 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5
Questions people ask
Which is better for workflow automation: GPT-5.3-Codex or Kimi K3?
GPT-5.3-Codex — it scored 90/100 against 85/100 on our 12-task workflow automation suite, winning 6 tasks to 1 with 5 tied. Every answer was marked blind by three judges from three rival AI labs.
How was this tested?
Both models answered the identical published workflow automation tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.
More workflow automation head-to-heads: GPT-5.3-Codex vs GPT-5.6 Sol · GPT-5.6 Sol vs Kimi K3 · GPT-5.3-Codex vs GPT-5.6 Terra · GPT-5.3-Codex vs GPT-5.6 Luna · Claude Sonnet 5 vs GPT-5.3-Codex · GPT-5.3-Codex vs GPT-5.5
Full ranking: Best AI for workflow automation · model pages: GPT-5.3-Codex, Kimi K3