Claude Opus 4.8 vs GPT-5.3-Codex for meeting notes

The verdict

GPT-5.3-Codex wins, 9188

Same 12 meeting notes tasks, marked blind by three rival labs. GPT-5.3-Codex took 4 tasks, Claude Opus 4.8 took 3, 5 tied. Tested 28 Aug 2026.

Where they differed most

The task — The transcript is wrong

Automatic transcription, and it has clearly misheard things. 'we need to get the S O W signed before we can start the disco very phase' … 'the client wants net thirty but our standard is net fifteen' … 'ARR is at four hundred and twenty K, sorry, four twenty thousand' … 'talk to Deloitte about the audit — no, not Deloitte, the other one, Grant Thornton'. Write the cleaned summary. Silently fix the transcription errors that are obvious from context, but where the AMBIGUITY IS REAL rather than a mishearing, flag it instead of guessing. One of these four is a genuine correction the speaker made out loud, not a transcription error, and it must be treated as such. Under 120 words.

**Summary**

- **SOW** (Statement of Work) needs to be signed before the discovery phase can begin.
- The client wants **Net 30** payment terms, but our standard is **Net 15**.
- **ARR is at $420K.** *(The "sorry, four twenty thousand" is the speaker's own live correction/restatement, not a mishearing — treating it as intended.)*
- Audit vendor: the speaker explicitly corrected themselves — **not Deloitte, but Grant Thornton.**

**Flag:** One point needs confirmation. The speaker's stated figures ("four hundred and twenty K" vs. "four twenty th…
We need the **SOW** signed before starting the **discovery phase**.  
The client wants **Net 30** payment terms, while our standard is **Net 15**.  
**ARR is $420,000**.  
For the audit, the speaker corrected themselves: **not Deloitte—Grant Thornton**.

Task by task

TaskClaude Opus 4.8GPT-5.3-Codex
The decision that was never made1010
Actions, owners, and the one with neither55
The number said twice810
Minutes for a meeting with a disagreement in it109
Fourteen minutes, one sentence108
Crosstalk109
The musing that is not a commitment1010
The follow-up email nobody has to rewrite1010
Extract the dates, resolve the relatives1010
Who has not spoken910
Notes for the person who missed it810
The transcript is wrong58

Full receipts: Claude Opus 4.8, GPT-5.3-Codex · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5

Questions people ask

Which is better for meeting notes: Claude Opus 4.8 or GPT-5.3-Codex?

GPT-5.3-Codex — it scored 91/100 against 88/100 on our 12-task meeting notes suite, winning 4 tasks to 3 with 5 tied. Every answer was marked blind by three judges from three rival AI labs.

How was this tested?

Both models answered the identical published meeting notes tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.

More meeting notes head-to-heads: GPT-5.3-Codex vs Grok 4.5 · Claude Opus 4.8 vs Grok 4.5 · GPT-5.3-Codex vs GPT-5.6 Sol · Claude Opus 4.8 vs GPT-5.6 Sol · GPT-5.3-Codex vs GPT-5.5 · Claude Opus 4.8 vs GPT-5.5

Full ranking: Best AI for meeting notes · model pages: Claude Opus 4.8, GPT-5.3-Codex