Claude Opus 4.8 vs GPT-5.6 Terra for bookkeeping
GPT-5.6 Terra wins, 93–83
Same 12 bookkeeping tasks, marked blind by three rival labs. GPT-5.6 Terra took 7 tasks, Claude Opus 4.8 took 3, 2 tied. Tested 18 Aug 2026.
Where they differed most
A bank reconciliation will not balance. The bank statement closing balance is £8,412.55. The cash book closing balance is £8,376.55. Outstanding items already accounted for: an unpresented cheque of £120.00 and a lodgement of £120.00 not yet on the statement, which cancel each other out. The bookkeeper has checked her additions twice. The cash book receipts for the month, as entered: £451.00, £1,208.30, £96.75, £2,340.00, £415.00, £188.50. The corresponding bank credits: £415.00, £1,208.30, £96.75, £2,340.00, £451.00, £188.50. Work out the difference, say what kind of error it is, explain the arithmetic rule that identifies this class of error at a glance, and say which side is wrong and how she would prove it. Then say what she should do about it in the ledger and what she should NOT do.
# Bank Reconciliation Error Analysis ## The Difference £8,412.55 − £8,376.55 = **£36.00** ## Finding the Error Comparing the cash book receipts with the bank credits: | Cash book | Bank statement | |-----------|----------------| | £451.00 | £415.00 | | £1,208.30 | £1,208.30 | | £96.75 | £96.75 | | £2,340.00 | £2,340.00 | | £415.00 | £451.00 | | £188.50 | £188.50 | The figures **£451.00** and **£415.00** appear in *reversed order*. The digits 5 and 1 have been swapped: £4**51** became £4**15**. ## Kind of Error This is a **transposition…
- Adjusted bank balance: £8,412.55 less unpresented cheque £120.00 plus lodgement not yet credited £120.00 = **£8,412.55** The two outstanding items net to nil. - Cash-book balance: **£8,376.55** - Difference: **£36.00**. The bank-side figure is £36.00 higher than the cash-book figure. The listed cash-book receipts total: \[ 451.00+1,208.30+96.75+2,340.00+415.00+188.50 =£4,699.55 \] The bank credits total exactly the same £4,699.55. They are merely in a different order. Therefore, those receipts do **not** explain the di…
Task by task
| Task | Claude Opus 4.8 | GPT-5.6 Terra |
|---|---|---|
| The client lunch and the sandwich | 9 | 8 |
| Categorise, and admit what you cannot | 9 | 10 |
| The invoice totals | 10 | 10 |
| Find the £36 | 3 | 8 |
| Profit but no money | 8 | 9 |
| Not your decision to make | 10 | 9 |
| The faded shoebox | 9 | 10 |
| It is my company, it is my money | 9 | 8 |
| The mileage claim | 10 | 10 |
| Capital or revenue | 8 | 10 |
| Quote, invoice, receipt | 9 | 10 |
| Refuse the backdated invoice | 5 | 10 |
Full receipts: Claude Opus 4.8, GPT-5.6 Terra · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5
Questions people ask
Which is better for bookkeeping: Claude Opus 4.8 or GPT-5.6 Terra?
GPT-5.6 Terra — it scored 93/100 against 83/100 on our 12-task bookkeeping suite, winning 7 tasks to 3 with 2 tied. Every answer was marked blind by three judges from three rival AI labs.
How was this tested?
Both models answered the identical published bookkeeping tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.
More bookkeeping head-to-heads: GPT-5.6 Sol vs GPT-5.6 Terra · GPT-5.3-Codex vs GPT-5.6 Terra · GPT-5.5 vs GPT-5.6 Terra · Claude Fable 5 vs GPT-5.6 Terra · GPT-5.6 Terra vs Grok 4.5 · GPT-5.6 Terra vs Qwen3.7 Max
Full ranking: Best AI for bookkeeping · model pages: Claude Opus 4.8, GPT-5.6 Terra