Compare AI models side by side
Not spec sheets — scores. Every model below sat the same tasks for each job, marked blind by three rival labs. Pick the ones you are choosing between.
| The job | GLM 5.2overall 92/100 API $0.966·$3.04 per 1M free via chat.z.ai | GPT-5.6 Terraoverall 92/100 API $2·$12 per 1M no verified free route | Grok 4.5overall 92/100 API $2·$6 per 1M free via grok.com |
|---|---|---|---|
| spreadsheets & Excel19 models tested | 96/100 · #6 | 98/100 · #2 | 94/100 · #8 |
| essay writing19 models tested | 79/100 · #14 | 88/100 · #4 | 83/100 · #8 |
| summarising documents19 models tested | 91/100 · #12 | 96/100 · #4 | 92/100 · #11 |
| extracting data from text19 models tested | 96/100 · #14 | 100/100 · #3 | 100/100 · #4 |
| writing emails19 models tested | 93/100 · #7 | 97/100 · #5 | 93/100 · #6 |
| everyday maths19 models tested | 97/100 · #10 | 99/100 · #2 | 98/100 · #5 |
| customer service replies19 models tested | 93/100 · #12 | 97/100 · #5 | 98/100 · #3 |
| revision & study19 models tested | 89/100 · #6 | 94/100 · #1 | 88/100 · #8 |
| coding20 models tested | 92/100 · #4 | 93/100 · #2 | 88/100 · #11 |
| vibe coding18 models tested | 81/100 · #13 | 95/100 · #6 | 90/100 · #7 |
| making flashcards19 models tested | 93/100 · #9 | 93/100 · #10 | 94/100 · #8 |
| social media posts19 models tested | 93/100 · #9 | 99/100 · #2 | 93/100 · #8 |
| job applications19 models tested | 95/100 · #2 | 89/100 · #9 | 88/100 · #12 |
| presentations and slides19 models tested | 87/100 · #12 | 96/100 · #4 | 91/100 · #9 |
| CV writing19 models tested | 85/100 · #13 | 91/100 · #6 | 88/100 · #9 |
| research skills19 models tested | 86/100 · #15 | 94/100 · #6 | 93/100 · #7 |
| creative writing19 models tested | 86/100 · #16 | 97/100 · #2 | 93/100 · #10 |
| translation18 models tested | 86/100 · #11 | 94/100 · #3 | 83/100 · #15 |
| travel planning19 models tested | 76/100 · #17 | 91/100 · #5 | 83/100 · #10 |
| emotional support19 models tested | 91/100 · #10 | 93/100 · #9 | 90/100 · #13 |
| everyday legal questions19 models tested | 86/100 · #10 | 93/100 · #3 | 86/100 · #11 |
| health questions19 models tested | 89/100 · #9 | 94/100 · #3 | 95/100 · #2 |
| book writing19 models tested | 90/100 · #8 | 96/100 · #3 | 90/100 · #9 |
| humanising AI text19 models tested | 93/100 · #6 | 93/100 · #7 | 90/100 · #10 |
| multi-step research19 models tested | 78/100 · #14 | 95/100 · #1 | 85/100 · #9 |
| code review18 models tested | 87/100 · #9 | 95/100 · #3 | 88/100 · #7 |
| workflow automation18 models tested | 82/100 · #13 | 90/100 · #3 | 86/100 · #7 |
| bookkeeping19 models tested | 81/100 · #11 | 93/100 · #1 | 87/100 · #6 |
| HR and employment questions18 models tested | 83/100 · #16 | 91/100 · #5 | 88/100 · #10 |
| property and lettings admin19 models tested | 78/100 · #15 | 95/100 · #1 | 83/100 · #10 |
| meeting notes19 models tested | 88/100 · #8 | 88/100 · #7 | 94/100 · #1 |
Every number is a score from the same published task suite for that job, marked blind by three rival labs, and every one links to the board it came from where the tasks and the models’ actual answers are printed in full. “#” is the model’s rank in that job’s full field, not just among the ones you picked — so #4 of 19 means something different from #4 of 5. “Declined” means the model refused the suite; that is its own answer, not a test we skipped.
Or read the head-to-head
Every pair we have a decisive result for, written up with both models’ real answers to each task. Pick the job, then the pair.
723 pairs with a decisive result. A further 672 finished close enough that the honest answer is “tied — pick on price”, so they are not listed here; the full ranking for each job shows every model including those.
spreadsheets & Excel
19 tested pairs · full ranking →
essay writing
33 tested pairs · full ranking →
summarising documents
21 tested pairs · full ranking →
writing emails
25 tested pairs · full ranking →
everyday maths
1 tested pairs · full ranking →
customer service replies
24 tested pairs · full ranking →
revision & study
25 tested pairs · full ranking →
coding
19 tested pairs · full ranking →
vibe coding
33 tested pairs · full ranking →
making flashcards
20 tested pairs · full ranking →
social media posts
24 tested pairs · full ranking →
job applications
30 tested pairs · full ranking →
presentations and slides
29 tested pairs · full ranking →
CV writing
24 tested pairs · full ranking →
research skills
28 tested pairs · full ranking →
creative writing
17 tested pairs · full ranking →
translation
25 tested pairs · full ranking →
travel planning
30 tested pairs · full ranking →
emotional support
18 tested pairs · full ranking →
everyday legal questions
18 tested pairs · full ranking →
health questions
18 tested pairs · full ranking →
book writing
30 tested pairs · full ranking →
humanising AI text
19 tested pairs · full ranking →
multi-step research
32 tested pairs · full ranking →
code review
27 tested pairs · full ranking →
workflow automation
27 tested pairs · full ranking →
bookkeeping
31 tested pairs · full ranking →
HR and employment questions
19 tested pairs · full ranking →
property and lettings admin
33 tested pairs · full ranking →
meeting notes
24 tested pairs · full ranking →