Claude Opus 4.6 vs Grok 4.5 for meeting notes
Grok 4.5 wins, 94–88
Same 12 meeting notes tasks, marked blind by three rival labs. Grok 4.5 took 3 tasks, Claude Opus 4.6 took 1, 8 tied. Tested 28 Aug 2026.
Where they differed most
Transcript extract, board update. FINANCE LEAD: 'Churn was four point two percent in July… sorry, four point two for the quarter, July on its own was three point one.' CHAIR: 'And that's up?' FINANCE LEAD: 'Up from two point eight the quarter before.' CHAIR: 'Right.' Write the three-line summary of the churn position for people who were not in the room. It must carry the corrected figures, must make clear which number is monthly and which is quarterly, and must not present the retracted reading as fact. Under 70 words.
**Churn Update** Quarterly churn rose to 4.2%, up from 2.8% in the prior quarter. Within that period, July's monthly churn stood at 3.1%. The Finance Lead initially cited the 4.2% figure as a July number before correcting it to a quarterly measure, so attendees should rely on the corrected breakdown above.
Churn was 4.2% for the quarter. This was up from 2.8% the prior quarter. July’s monthly churn was 3.1%.
Task by task
| Task | Claude Opus 4.6 | Grok 4.5 |
|---|---|---|
| The decision that was never made | 10 | 10 |
| Actions, owners, and the one with neither | 10 | 10 |
| The number said twice | 5 | 10 |
| Minutes for a meeting with a disagreement in it | 9 | 10 |
| Fourteen minutes, one sentence | 10 | 10 |
| Crosstalk | 10 | 10 |
| The musing that is not a commitment | 10 | 10 |
| The follow-up email nobody has to rewrite | 10 | 10 |
| Extract the dates, resolve the relatives | 10 | 10 |
| Who has not spoken | 10 | 8 |
| Notes for the person who missed it | 6 | 6 |
| The transcript is wrong | 5 | 9 |
Full receipts: Claude Opus 4.6, Grok 4.5 · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5
Questions people ask
Which is better for meeting notes: Claude Opus 4.6 or Grok 4.5?
Grok 4.5 — it scored 94/100 against 88/100 on our 12-task meeting notes suite, winning 3 tasks to 1 with 8 tied. Every answer was marked blind by three judges from three rival AI labs.
How was this tested?
Both models answered the identical published meeting notes tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.
More meeting notes head-to-heads: GPT-5.6 Sol vs Grok 4.5 · GPT-5.5 vs Grok 4.5 · GPT-5.3-Codex vs Grok 4.5 · Grok 4.5 vs Qwen3.7 Max · GPT-5.6 Terra vs Grok 4.5 · GLM 5.2 vs Grok 4.5
Full ranking: Best AI for meeting notes · model pages: Claude Opus 4.6, Grok 4.5