GPT-5.6 Sol vs Grok 4.5 for bookkeeping
GPT-5.6 Sol wins, 93–87
Same 12 bookkeeping tasks, marked blind by three rival labs. GPT-5.6 Sol took 7 tasks, Grok 4.5 took 1, 4 tied. Tested 18 Aug 2026.
Where they differed most
Guidance excerpt: 'Capital expenditure is money spent acquiring, creating or improving an asset that will be used in the business over more than one accounting period. Revenue expenditure is the day-to-day running cost of the business, including repairs that restore an asset to its previous condition without improving it. Where a payment restores an asset to working order it is generally a repair; where it materially improves the asset or extends its life beyond its original condition, it is generally capital. Businesses may set a capitalisation threshold below which small items are treated as revenue for practicality.' Classify these six, using only the excerpt. Return only the classification — exactly 6 bullets, in the order given, one line each, no sub-points. Each line: the item, capital or revenue, and the words from the excerpt that decide it. Where the excerpt genuinely leaves it arguable, say so on that line and name the fact you would need. 1) New laptop, £1,100. 2) Repair to the same laptop's screen, £180. 3) New shelving racking for the storeroom, £340. 4) Annual accounting software subscription, £480. 5) Replacement engine for the works van, £2,200. 6) Van MOT and service, £54.
- New laptop, £1,100 — capital, unless below the business’s capitalisation threshold; “acquiring…an asset that will be used…over more than one accounting period”; need the threshold. - Repair to the same laptop’s screen, £180 — revenue; “repairs that restore an asset to its previous condition without improving it.” - New shelving racking for the storeroom, £340 — capital, unless below the business’s capitalisation threshold; “acquiring…an asset that will be used…over more than one accounting period”; need the threshold. - Annual accounting soft…
- 1) New laptop, £1,100: capital - "money spent acquiring... an asset that will be used in the business over more than one accounting period" - 2) Repair to the same laptop's screen, £180: revenue - "repairs that restore an asset to its previous condition without improving it" / "restores an asset to working order it is generally a repair" - 3) New shelving racking for the storeroom, £340: capital - "money spent acquiring... an asset that will be used in the business over more than one accounting period" (arguable if below capitalisation thresh…
Task by task
| Task | GPT-5.6 Sol | Grok 4.5 |
|---|---|---|
| The client lunch and the sandwich | 8 | 8 |
| Categorise, and admit what you cannot | 9 | 8 |
| The invoice totals | 10 | 10 |
| Find the £36 | 9 | 8 |
| Profit but no money | 8 | 10 |
| Not your decision to make | 10 | 9 |
| The faded shoebox | 10 | 9 |
| It is my company, it is my money | 8 | 7.5 |
| The mileage claim | 10 | 10 |
| Capital or revenue | 10 | 7 |
| Quote, invoice, receipt | 9 | 8 |
| Refuse the backdated invoice | 10 | 10 |
Full receipts: GPT-5.6 Sol, Grok 4.5 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5
Questions people ask
Which is better for bookkeeping: GPT-5.6 Sol or Grok 4.5?
GPT-5.6 Sol — it scored 93/100 against 87/100 on our 12-task bookkeeping suite, winning 7 tasks to 1 with 4 tied. Every answer was marked blind by three judges from three rival AI labs.
How was this tested?
Both models answered the identical published bookkeeping tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.
More bookkeeping head-to-heads: GPT-5.5 vs GPT-5.6 Sol · GPT-5.6 Sol vs Kimi K3 · Claude Opus 4.8 vs GPT-5.6 Sol · GLM 5.2 vs GPT-5.6 Sol · Claude Sonnet 5 vs GPT-5.6 Sol · Gemini 3.5 Flash vs GPT-5.6 Sol
Full ranking: Best AI for bookkeeping · model pages: GPT-5.6 Sol, Grok 4.5