Best AI for bookkeeping

The verdict

GPT-5.6 Sol

Scored 93/100 on our 12-task bookkeeping suitea single point ahead of GPT-5.5 (92) — effectively level.

What to actually do

GPT-5.6 Sol is the engine inside ChatGPT. Go to chatgpt.comthe free tier is fine to start. Paid plans start at £7/month (go plan, vendor’s own price). The free tier is ChatGPT’s, not a promise about this exact model — we haven’t verified which plan carries it. Not fussed about the last point or two? Any of the top 2 here will serve you well.

UK small-business bookkeeping, and the rule is always quoted in the task: answers are marked against the guidance in front of them, never against tax law the model claims to remember. A client lunch and a motorway sandwich, nine lines off a bank feed with the ones that cannot be categorised flagged rather than guessed, a reconciliation £36 out for a reason a bookkeeper spots in seconds, and an invoice whose four totals have exactly one correct answer, compared by the harness rather than by opinion. One task is the call every bookkeeper gets — £4,180 profit, £900 in the bank — answered in 120 words without using the word 'debtors'. None of this is accountancy or tax advice — answers are judged against the guidance quoted in the task, and a good one tells you when to ring your accountant. One task asks for an invoice to be redated across the year end and a daughter who does not work there put on the payroll; the right answer is to refuse both, and to say which question belongs with the accountant.

updated 18 Aug 2026 · tested by Robert Prime · re-ranks automatically when a new run lands

#ModelOur score
1GPT-5.6 Sollatest93/100
2GPT-5.592/100
3Grok 4.5latest87/100
4Kimi K384/100
5Claude Opus 4.883/100
6GLM 5.281/100
7Claude Sonnet 581/100
8Gemini 3.5 Flash81/100
9Gemini 3.1 Pro Preview78/100
10DeepSeek V4 Pro70/100

“API cost” is what software developers pay to build on a model — ignore it if you just use the website. Each model answers each task once. Models level on score are ranked by a fixed tie-break — fewest machine-checked rule breaches, then lowest measured cost per run — so the order is deterministic and checkable, never arbitrary. Judge panels never include the contestant’s own lab, so panels differ slightly per model — small cross-model gaps can reflect panel severity, not quality.

1.

GPT-5.6 Sol

93/100our pick — won on the tie-break

Made by OpenAI — their newest model. You use it inside ChatGPT — nothing to install.

Strongest showing: Find the £36” — scored 9/10 by the panel. Weakest: “Profit but no money” at 8/10. On one task a judge flagged a possible instruction breach — advisory only: the machine check did not confirm it and the panel outvoted it.

Accurate difference, transposition diagnosis via /9 rule, correct conclusion neither list errs, proper proof/ledger advice. Minor polish only.x-ai/grok-4.5, judging blind · full receipts ↓

12 tasks · 18 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $2.5 in / $15 out per 1M tokens · full model page →

2.

GPT-5.5

92/100

Made by OpenAI. You use it inside ChatGPT — nothing to install.

Strongest showing: Find the £36” — scored 8/10 by the panel.

Correct diff and transposition identification, good rule explanation, sound proof/correction advice. Minor imprecision: journal entry direction slightly informal, and 'divisible by 9' generalization stated loosely. Overall clear, useful, acanthropic/claude-sonnet-5, judging blind · full receipts ↓

12 tasks · 18 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $5 in / $30 out per 1M tokens · full model page →

3.

Grok 4.5

87/100

Made by xAI — their newest model. You use it inside Grok — nothing to install; free via grok.com.

Strongest showing: Find the £36” — scored 8/10 by the panel. Weakest: “Capital or revenue” at 7/10. On 2 tasks a judge flagged a possible instruction breach — advisory only: the machine check did not confirm it and the panel outvoted it.

Correct diff (£36), correct transposition error identification, correct divisibility-by-9 rule, clear proof method and ledger guidance; concise and accurate overall.anthropic/claude-sonnet-5, judging blind · full receipts ↓

12 tasks · 18 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $2 in / $6 out per 1M tokens · free via grok.com (checked 11 Aug 2026) · full model page →

See it for yourself

One task from the suite, and what the leader, the middle of the table and the bottom actually wrote. This is the task they disagreed on most — scores ranged from 3 to 9 out of 10.

The task — Find the £36

A bank reconciliation will not balance. The bank statement closing balance is £8,412.55. The cash book closing balance is £8,376.55. Outstanding items already accounted for: an unpresented cheque of £120.00 and a lodgement of £120.00 not yet on the statement, which cancel each other out. The bookkeeper has checked her additions twice. The cash book receipts for the month, as entered: £451.00, £1,208.30, £96.75, £2,340.00, £415.00, £188.50. The corresponding bank credits: £415.00, £1,208.30, £96.75, £2,340.00, £451.00, £188.50. Work out the difference, say what kind of error it is, explain the arithmetic rule that identifies this class of error at a glance, and say which side is wrong and how she would prove it. Then say what she should do about it in the ledger and what she should NOT do.

- **Adjusted bank-statement balance:**  
  £8,412.55 − £120 unpresented cheque + £120 outstanding lodgement = **£8,412.55**

- **Difference:**  
  £8,412.55 − £8,376.55 = **£36.00**  
  Thus, if the statement is correct, the cash-book balance is £36 too low.

### Apparent type of error

£36 has the…

Judge: Accurate difference, transposition diagnosis via /9 rule, correct conclusion neither list errs, proper proof/ledger advice. Minor polish only.

**1. The Difference**
The difference between the bank statement closing balance (£8,412.55) and the cash book closing balance (£8,376.55) is **£36.00**. 

**2. Kind of Error**
This is a **transposition error**. Specifically, the figures £451.00 and £415.00 have been transposed. The bookkeeper has en…

Judge: The response incorrectly identifies a transposition error between £451 and £415 in the provided lists, failing to notice that both amounts appear in both the cash book and bank statement lists. This is a major factual error.

The difference is **£36.00** (£8,412.55 – £8,376.55).  

**Kind of error:** Transposition error (digits swapped).  

**Arithmetic rule:** When the difference between two numbers is exactly divisible by 9, it is a strong indicator of a transposition error. Here, £36 ÷ 9 = 4, pointing to a swap of dig…

Judge: The response incorrectly claims that entering £451 and £415 in the opposite order causes a £36 difference. Since both amounts appear on both sides, their sum is identical, so this cannot explain the £36 discrepancy.

Trimmed to fit. Every full answer, every judge’s score: GPT-5.6 Sol, GLM 5.2, DeepSeek V4 Pro

How this ranking is made

Every model answers the identical published task suite. Three AI judges from three different labs mark each answer blind against a fixed 0–10 rubric — none is ever from the maker’s own lab, and an answer that breaks a task’s explicit rules is capped in code, not by the judge’s goodwill. The score is the average, out of 100.

What this is not: an opinion piece, a paid ranking, or a benchmark we cannot show you. No vendor can buy inclusion, a position or a score on this page — the order is computed from the test results before any link to a product exists, and where a link earns us a commission it says so on the link itself and the order is identical either way. Every score links its raw outputs and judge verdicts. The full protocol · How we make money · receipts: GPT-5.6 Sol, GPT-5.5, Grok 4.5, Kimi K3, Claude Opus 4.8, GLM 5.2, Claude Sonnet 5, Gemini 3.5 Flash, Gemini 3.1 Pro Preview, DeepSeek V4 Pro

Questions people ask

What is the best AI for bookkeeping in 2026?

GPT-5.6 Sol leads our tested ranking with 93/100 on our 12-task bookkeeping suite (12 tasks), in a dead heat with GPT-5.5 (92). Every answer was marked blind by three AI judges from three different labs, and the full outputs are downloadable.

How is this ranking made?

Each model answers the identical published task suite; three judges from different labs score every answer 0–10 against a fixed rubric without knowing which produced it; answers that break a task's explicit rules are capped automatically. The score is the average, out of 100. No vendor pays for placement.

How often does this page update?

It re-ranks itself whenever a new test run lands, and prices re-verify daily against vendor pages. The current ranking was last computed on 18 Aug 2026.

Head-to-heads in bookkeeping & accounts

Show all 20 tested pairs ▾

All comparisons →

More rankings ▾

Best AI for writing · Best AI chatbot for everyday use · Best AI for coding · Best free AI model · Best AI for spreadsheets and Excel · Best AI essay writer · Best AI for summarising documents · Best AI for extracting data from text · Best AI for writing emails · Best AI for everyday maths and percentages · Best AI for customer service replies · Best AI for revision and study notes · Best AI for vibe coding · Best AI for making flashcards · Best AI for social media posts · Best AI for job applications and cover letters · Best AI for presentations · Best AI for writing your CV · Best AI for research skills · Best AI for creative writing · Best AI for translation · Best AI for travel planning · Best AI for emotional support · Best AI for everyday legal questions · Best AI for health questions · Best AI for writing a book · Best AI for humanising AI text · Best AI research agent · Best AI for code review · Best AI for workflow automation · Best AI for HR and employment questions · Best AI for property and lettings · Best-value AI model API · every model we track · every tool