DeepSeek V4 Pro vs GPT-5.3-Codex for presentations and slides
GPT-5.3-Codex wins, 97–91
Same 12 presentations and slides tasks, marked blind by three rival labs. GPT-5.3-Codex took 5 tasks, DeepSeek V4 Pro took 2, 5 tied. Tested 16 Aug 2026.
Where they differed most
You are presenting a proposal to cut free staff parking to fund an extra day of annual leave for everyone. Write the 5 hardest questions the audience will actually ask, numbered 1-5, one line each with the honest one-sentence answer on the same line after a dash. No sub-points, no softballs.
1. How will you mitigate the inequity for low-paid staff who can't absorb new daily parking costs? – We'll offer a salary-equivalent opt-out cash allowance for non-parkers, making the leave benefit tradable and redirecting the net savings to a hardship fund that offsets the precise shortfall for the lowest quartile of earners. 2. Won't this simply shift parking demand onto surrounding residential streets and trigger enforcement wars with neighbors? – We've negotiated a paid permit scheme with the council using a portion of the savings, and th…
1. Why should staff who have no realistic alternative to driving (shift workers, carers, disabled colleagues) effectively pay for everyone else’s extra leave? — They shouldn’t, so the proposal only works if we ring‑fence exemptions or subsidies for genuinely car-dependent staff and accept a smaller net saving. 2. Are you certain the parking cuts actually cover the full cost of one extra leave day, including backfill/overtime in frontline teams? — No, not with certainty, so we need transparent modelling with a contingency because covering abse…
Task by task
| Task | DeepSeek V4 Pro | GPT-5.3-Codex |
|---|---|---|
| Deck from a standing start | 8 | 9 |
| Speaker notes, not an essay | 8 | 10 |
| One slide, not a memo | 10 | 9 |
| The first 30 seconds | 10 | 10 |
| Slide on a diet | 10 | 10 |
| Right chart, right headline | 10 | 10 |
| Hostile Q&A prep | 6 | 10 |
| Cut it to two minutes | 9 | 10 |
| Same facts, different room | 10 | 9 |
| Forty slides into five | 10 | 10 |
| Father of the bride | 8 | 9 |
| Refuse the inflated slide | 10 | 10 |
Full receipts: DeepSeek V4 Pro, GPT-5.3-Codex · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra
Questions people ask
Which is better for presentations and slides: DeepSeek V4 Pro or GPT-5.3-Codex?
GPT-5.3-Codex — it scored 97/100 against 91/100 on our 12-task presentations and slides suite, winning 5 tasks to 2 with 5 tied. Every answer was marked blind by three judges from three rival AI labs.
How was this tested?
Both models answered the identical published presentations and slides tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.
More presentations and slides head-to-heads: GPT-5.3-Codex vs GPT-5.6 Luna · DeepSeek V4 Pro vs GPT-5.6 Luna · GPT-5.3-Codex vs GPT-5.5 · DeepSeek V4 Pro vs GPT-5.5 · GPT-5.3-Codex vs GPT-5.6 Terra · GPT-5.3-Codex vs GPT-5.6 Sol
Full ranking: Best AI for presentations and slides · model pages: DeepSeek V4 Pro, GPT-5.3-Codex