Compare AI models side by side

Not spec sheets — scores. Every model below sat the same tasks for each job, marked blind by three rival labs. Pick the ones you are choosing between.

Pick the models to compare — 3 selected
Sorting by where these models actually disagree. Jobs where they all score the same are not a reason to choose one.
The jobGLM 5.2overall 92/100
API $0.966·$3.04 per 1M
free via chat.z.ai
GPT-5.6 Terraoverall 92/100
API $2·$12 per 1M
no verified free route
Grok 4.5overall 92/100
API $2·$6 per 1M
free via grok.com
spreadsheets & Excel19 models tested96/100 · #698/100 · #294/100 · #8
essay writing19 models tested79/100 · #1488/100 · #483/100 · #8
summarising documents19 models tested91/100 · #1296/100 · #492/100 · #11
extracting data from text19 models tested96/100 · #14100/100 · #3100/100 · #4
writing emails19 models tested93/100 · #797/100 · #593/100 · #6
everyday maths19 models tested97/100 · #1099/100 · #298/100 · #5
customer service replies19 models tested93/100 · #1297/100 · #598/100 · #3
revision & study19 models tested89/100 · #694/100 · #188/100 · #8
coding20 models tested92/100 · #493/100 · #288/100 · #11
vibe coding18 models tested81/100 · #1395/100 · #690/100 · #7
making flashcards19 models tested93/100 · #993/100 · #1094/100 · #8
social media posts19 models tested93/100 · #999/100 · #293/100 · #8
job applications19 models tested95/100 · #289/100 · #988/100 · #12
presentations and slides19 models tested87/100 · #1296/100 · #491/100 · #9
CV writing19 models tested85/100 · #1391/100 · #688/100 · #9
research skills19 models tested86/100 · #1594/100 · #693/100 · #7
creative writing19 models tested86/100 · #1697/100 · #293/100 · #10
translation18 models tested86/100 · #1194/100 · #383/100 · #15
travel planning19 models tested76/100 · #1791/100 · #583/100 · #10
emotional support19 models tested91/100 · #1093/100 · #990/100 · #13
everyday legal questions19 models tested86/100 · #1093/100 · #386/100 · #11
health questions19 models tested89/100 · #994/100 · #395/100 · #2
book writing19 models tested90/100 · #896/100 · #390/100 · #9
humanising AI text19 models tested93/100 · #693/100 · #790/100 · #10
multi-step research19 models tested78/100 · #1495/100 · #185/100 · #9
code review18 models tested87/100 · #995/100 · #388/100 · #7
workflow automation18 models tested82/100 · #1390/100 · #386/100 · #7
bookkeeping19 models tested81/100 · #1193/100 · #187/100 · #6
HR and employment questions18 models tested83/100 · #1691/100 · #588/100 · #10
property and lettings admin19 models tested78/100 · #1595/100 · #183/100 · #10
meeting notes19 models tested88/100 · #888/100 · #794/100 · #1

Every number is a score from the same published task suite for that job, marked blind by three rival labs, and every one links to the board it came from where the tasks and the models’ actual answers are printed in full. “#” is the model’s rank in that job’s full field, not just among the ones you picked — so #4 of 19 means something different from #4 of 5. “Declined” means the model refused the suite; that is its own answer, not a test we skipped.

Or read the head-to-head

Every pair we have a decisive result for, written up with both models’ real answers to each task. Pick the job, then the pair.

723 pairs with a decisive result. A further 672 finished close enough that the honest answer is “tied — pick on price”, so they are not listed here; the full ranking for each job shows every model including those.

spreadsheets & Excel

19 tested pairs · full ranking →

essay writing

33 tested pairs · full ranking →

summarising documents

21 tested pairs · full ranking →

writing emails

25 tested pairs · full ranking →

everyday maths

1 tested pairs · full ranking →

customer service replies

24 tested pairs · full ranking →

revision & study

25 tested pairs · full ranking →

coding

19 tested pairs · full ranking →

vibe coding

33 tested pairs · full ranking →

making flashcards

20 tested pairs · full ranking →

social media posts

24 tested pairs · full ranking →

job applications

30 tested pairs · full ranking →

presentations and slides

29 tested pairs · full ranking →

CV writing

24 tested pairs · full ranking →

research skills

28 tested pairs · full ranking →

creative writing

17 tested pairs · full ranking →

translation

25 tested pairs · full ranking →

travel planning

30 tested pairs · full ranking →

emotional support

18 tested pairs · full ranking →

everyday legal questions

18 tested pairs · full ranking →

health questions

18 tested pairs · full ranking →

book writing

30 tested pairs · full ranking →

humanising AI text

19 tested pairs · full ranking →

multi-step research

32 tested pairs · full ranking →

code review

27 tested pairs · full ranking →

workflow automation

27 tested pairs · full ranking →

bookkeeping

31 tested pairs · full ranking →

HR and employment questions

19 tested pairs · full ranking →

property and lettings admin

33 tested pairs · full ranking →

meeting notes

24 tested pairs · full ranking →