GPT-5.3-Codex vs GPT-5.6 Sol for meeting notes
Effectively a dead heat: 93–91
Same 12 meeting notes tasks, marked blind by three rival labs. GPT-5.6 Sol took 4 tasks, GPT-5.3-Codex took 1, 7 tied. Tested 28 Aug 2026.
Where they differed most
Two people disagreed, publicly, and were not reconciled. RAJ: 'Moving support to the outsourcer will cost us the thing customers actually rate us for.' KATE: 'It costs us eleven grand a month we do not have.' RAJ: 'And churn costs more.' KATE: 'You don't know that.' RAJ: 'Neither do you.' CHAIR: 'We'll take it to the September board with numbers on both.' Write formal minutes of this item. Record both positions with equal weight and named attribution, record the chair's ruling, and record that the disagreement is unresolved. Do not smooth it into consensus, do not editorialise about who is right, and do not use the words 'productive', 'robust' or 'healthy' to describe the exchange. Under 150 words.
**Item: Proposal to outsource support function** Raj stated that moving support to the outsourcer would risk the service element customers currently rate the company for, and argued that any resulting churn could exceed current costs. Kate stated that retaining support in-house currently costs approximately £11,000 per month, which she said the company cannot afford, and challenged whether churn impacts had been quantified. Raj responded that churn costs were material; Kate responded that this had not been demonstrated; Raj noted that neithe…
**Outsourcing of Customer Support** Raj opposed moving customer support to the outsourcer, stating that it risked undermining the service customers valued and that resulting customer churn could cost more. Kate supported the move on financial grounds, stating that the current arrangement cost 11,000 per month and that funds were unavailable to sustain it. She disputed that the cost of churn would be greater, noting that this had not been established. Raj responded that the financial case for outsourcing was likewise unproven. The Chair rule…
Task by task
| Task | GPT-5.3-Codex | GPT-5.6 Sol |
|---|---|---|
| The decision that was never made | 10 | 10 |
| Actions, owners, and the one with neither | 5 | 5 |
| The number said twice | 10 | 10 |
| Minutes for a meeting with a disagreement in it | 9 | 10 |
| Fourteen minutes, one sentence | 8 | 9 |
| Crosstalk | 9 | 10 |
| The musing that is not a commitment | 10 | 10 |
| The follow-up email nobody has to rewrite | 10 | 10 |
| Extract the dates, resolve the relatives | 10 | 10 |
| Who has not spoken | 10 | 10 |
| Notes for the person who missed it | 10 | 9 |
| The transcript is wrong | 8 | 9 |
Full receipts: GPT-5.3-Codex, GPT-5.6 Sol · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5
Questions people ask
Which is better for meeting notes: GPT-5.3-Codex or GPT-5.6 Sol?
Effectively a dead heat: GPT-5.6 Sol edged it 93/100 to 91/100 on our meeting notes suite — too close to matter, so pick on price or the product you already use.
How was this tested?
Both models answered the identical published meeting notes tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.
More meeting notes head-to-heads: GPT-5.6 Sol vs Grok 4.5 · GPT-5.3-Codex vs Grok 4.5 · GPT-5.5 vs GPT-5.6 Sol · GPT-5.6 Sol vs Qwen3.7 Max · Claude Opus 4.6 vs GPT-5.6 Sol · GPT-5.6 Sol vs GPT-5.6 Terra
Full ranking: Best AI for meeting notes · model pages: GPT-5.3-Codex, GPT-5.6 Sol