GPT-5.6 Terra vs Grok 4.5 for meeting notes
Grok 4.5 wins, 94–88
Same 12 meeting notes tasks, marked blind by three rival labs. Grok 4.5 took 4 tasks, GPT-5.6 Terra took 2, 6 tied. Tested 28 Aug 2026.
Where they differed most
Transcript extract. MAYA: 'I'll get the pricing page updated before the campaign goes out.' TOM: 'And someone needs to tell support about the new tiers.' MAYA: 'Yeah.' TOM: 'I can do the deck for Thursday.' MAYA: 'Perfect. Oh and the legal review — Ash said he'd look but that was three weeks ago.' TOM: 'I'll chase Ash.' MAYA: 'Actually leave it, I'll chase him, I owe him an email anyway.' Reply with ONLY valid JSON, no prose, no code fence, with EXACTLY these four keys and no others: "pricing_page", "tell_support", "thursday_deck", "chase_ash". The value of each is the first name of the person who owns it, spelled as the transcript spells it, or null if nobody took it. Where ownership changed during the conversation, give the FINAL owner.
{"pricing_page":"MAYA","tell_support":null,"thursday_deck":"TOM","chase_ash":"MAYA"}{"pricing_page": "Maya", "tell_support": null, "thursday_deck": "Tom", "chase_ash": "Maya"}Task by task
| Task | GPT-5.6 Terra | Grok 4.5 |
|---|---|---|
| The decision that was never made | 10 | 10 |
| Actions, owners, and the one with neither | 5 | 10 |
| The number said twice | 10 | 10 |
| Minutes for a meeting with a disagreement in it | 10 | 10 |
| Fourteen minutes, one sentence | 9 | 10 |
| Crosstalk | 10 | 10 |
| The musing that is not a commitment | 10 | 10 |
| The follow-up email nobody has to rewrite | 9 | 10 |
| Extract the dates, resolve the relatives | 10 | 10 |
| Who has not spoken | 10 | 8 |
| Notes for the person who missed it | 7 | 6 |
| The transcript is wrong | 6 | 9 |
Full receipts: GPT-5.6 Terra, Grok 4.5 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5
Questions people ask
Which is better for meeting notes: GPT-5.6 Terra or Grok 4.5?
Grok 4.5 — it scored 94/100 against 88/100 on our 12-task meeting notes suite, winning 4 tasks to 2 with 6 tied. Every answer was marked blind by three judges from three rival AI labs.
How was this tested?
Both models answered the identical published meeting notes tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.
More meeting notes head-to-heads: GPT-5.6 Sol vs Grok 4.5 · GPT-5.5 vs Grok 4.5 · GPT-5.3-Codex vs Grok 4.5 · Grok 4.5 vs Qwen3.7 Max · Claude Opus 4.6 vs Grok 4.5 · GLM 5.2 vs Grok 4.5
Full ranking: Best AI for meeting notes · model pages: GPT-5.6 Terra, Grok 4.5