GPT-5.3-Codex vs GPT-5.5 for meeting notes
Effectively a dead heat: 92–91
Same 12 meeting notes tasks, marked blind by three rival labs. GPT-5.5 took 3 tasks, GPT-5.3-Codex took 2, 7 tied. Tested 28 Aug 2026.
Where they differed most
A fourteen-minute agenda item, summarised: the team debated whether to rebuild the onboarding flow or patch it; engineering estimated six weeks for a rebuild and 'about a fortnight, probably' for patching; design argued the patch preserves the confusing three-step signup that generates most support tickets; the head of product asked what the support-ticket volume actually is and nobody in the room knew; the item ended with someone agreeing to pull the ticket numbers before any decision. Write ONE sentence, maximum 30 words, that a director skimming the notes would read instead of the fourteen minutes. It must land the actual state of play — that the decision is blocked on a number nobody had.
Decision deferred: choosing between a six-week onboarding rebuild or ~fortnight patch is blocked until support-ticket volume for the confusing three-step signup is pulled, which nobody in the room had.
Decision on rebuilding versus patching onboarding is deferred until support-ticket data is pulled, since the key volume driving the tradeoff was unknown.
Task by task
| Task | GPT-5.3-Codex | GPT-5.5 |
|---|---|---|
| The decision that was never made | 10 | 10 |
| Actions, owners, and the one with neither | 5 | 5 |
| The number said twice | 10 | 10 |
| Minutes for a meeting with a disagreement in it | 9 | 9 |
| Fourteen minutes, one sentence | 8 | 10 |
| Crosstalk | 9 | 10 |
| The musing that is not a commitment | 10 | 10 |
| The follow-up email nobody has to rewrite | 10 | 10 |
| Extract the dates, resolve the relatives | 10 | 10 |
| Who has not spoken | 10 | 9 |
| Notes for the person who missed it | 10 | 8 |
| The transcript is wrong | 8 | 9 |
Full receipts: GPT-5.3-Codex, GPT-5.5 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5
Questions people ask
Which is better for meeting notes: GPT-5.3-Codex or GPT-5.5?
Effectively a dead heat: GPT-5.5 edged it 92/100 to 91/100 on our meeting notes suite — too close to matter, so pick on price or the product you already use.
How was this tested?
Both models answered the identical published meeting notes tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.
More meeting notes head-to-heads: GPT-5.5 vs Grok 4.5 · GPT-5.3-Codex vs Grok 4.5 · GPT-5.5 vs GPT-5.6 Sol · GPT-5.3-Codex vs GPT-5.6 Sol · GPT-5.5 vs Qwen3.7 Max · Claude Opus 4.6 vs GPT-5.5
Full ranking: Best AI for meeting notes · model pages: GPT-5.3-Codex, GPT-5.5