Which AI is best at which job?
We give every model the same set of real tasks and have three rival AI companies mark the answers blind. Here is who won what — and you can read every answer they gave.
473 ranked results, distilled from 524 scored test runs · no vendor pays to be here · updated 17 Aug 2026
Writing & communicating
Everyday writingdocuments, tone, rewrites90/100ChatGPTfree to start · paid from £7/moEmails & letterstone, brevity, the difficult ones99/100ChatGPT — two of its versions tiedruns the GPT-5.6 Sol model · free to start · paid from £7/moEssays & long-formarguments, word counts, structure95/100ChatGPTruns the GPT-5.6 Sol model · free to start · paid from £7/moCustomer repliescomplaints, refunds, reviews100/100ChatGPTruns the GPT-5.6 Sol model · free to start · paid from £7/moSocial postsLinkedIn & X, minus the cringe100/100ChatGPTruns the GPT-5.6 Sol model · free to start · paid from £7/moJob applicationscover letters, CV bullets98/100ChatGPTruns the GPT-5.5 model · free to start · paid from £7/moCV writingprofiles, bullets, the honest line95/100ChatGPTruns the GPT-5.5 model · free to start · paid from £7/moPresentationsslides, talk tracks, openings98/100ChatGPT — two of its versions tiedruns the GPT-5.6 Luna model · free to start · paid from £7/moTranslationFrench, Spanish, German — register intact96/100ChatGPTruns the GPT-5.6 Sol model · free to start · paid from £7/moHumanising AI textAI-sounding drafts made human, facts kept97/100ChatGPTruns the GPT-5.3-Codex model · free to start · paid from £7/mo
Working with information
Summarisinglong or messy input, short output99/100Clauderuns the Claude Opus 4.8 model · free to start · paid from $20/mo (US price)Pulling out datatext into fields, JSON, tables100/100Gemini · dead heat with 9 othersruns the Gemini 3.1 Flash Lite model · paid from £7.99/moSpreadsheets & Excelformulas, cleaning, checking numbers98/100ChatGPT — three of its versions tiedruns the GPT-5.6 Luna model · free to start · paid from £7/moEveryday mathsVAT, discounts, percentages100/100Clauderuns the Claude Opus 4.8 model · free to start · paid from $20/mo (US price)Research skillssources, evidence, answerable questions98/100DeepSeek · ties GPT-5.6 Solruns the DeepSeek V4 Pro model · at chat.deepseek.com
Learning & studying
Revision & studynotes, exam plans, explanations94/100ChatGPT — two of its versions tiedruns the GPT-5.6 Terra model · free to start · paid from £7/moFlashcardstestable cards from your notes98/100ChatGPT · ties Qwen3.7 Maxruns the GPT-5.6 Sol model · free to start · paid from £7/moEveryday questionsexplaining, planning, general help90/100ChatGPTfree to start · paid from £7/mo
Creative & books
Life & big questions
Travel planningmessy constraints into a real plan94/100ChatGPTruns the GPT-5.6 Sol model · free to start · paid from £7/moEmotional supportthe right words, and the right signposts98/100Claude · ties GPT-5.5runs the Claude Sonnet 5 model · free to start · paid from $20/mo (US price)Legal questionstenancy, refunds, parking — plain English94/100Claude · ties GPT-5.6 Solruns the Claude Opus 4.8 model · free to start · paid from $20/mo (US price)Health questionsmarked against public guidance95/100DeepSeek · ties Grok 4.5runs the DeepSeek V4 Pro model · at chat.deepseek.com
Building & coding
Where to spend your money
What to use if the bill matters more than the last point of quality.
Every score comes with the raw answers and the judges’ reasoning — here is how the testing works. Choosing a product rather than a model? What each tool costs and does → Torn between two specific ones? Every head-to-head we’ve tested →