Best AI for customer service replies

The verdict

GPT-5.6 Sol

Scored 100/100 on our 12-task customer-reply suitea single point ahead of GPT-5.5 (99) — effectively level.

What to actually do

GPT-5.6 Sol is the engine inside ChatGPT. Go to chatgpt.comthe free tier is fine to start. Paid plans start at £7/month (go plan, vendor’s own price). The free tier is ChatGPT’s, not a promise about this exact model — we haven’t verified which plan carries it. Not fussed about the last point or two? Any of the top 2 here will serve you well.

The hardest writing in business is to a customer who is angry, confused or right. This suite tests the replies that decide whether they stay: the instant refund, the clear no, the public 1-star response, the mid-outage update with no fix yet. Banned phrases are enforced — no 'unfortunately', no 'we take this seriously'. One task asks for five fake 5-star reviews; the right answer is to decline.

updated 14 Aug 2026 · tested by Robert Prime · re-ranks automatically when a new run lands

How this was measured
  • All 19 sat the identical 12-task suite — same tasks, same order, one attempt each.
  • Every answer marked blind: judges are not told which entrant wrote it.
  • Three judges per answer, each from a competing lab. A panel never includes the entrant's own lab.
  • Scored 0–10 against a fixed rubric. Answers breaking a task's explicit rules are capped by machine, not by opinion.
  • Ties broken by fewest rule breaches, then lowest measured cost — published, not editorial.
  • Nobody pays for placement. Last computed 14 Aug 2026.

Every score below links to the raw file behind it: every task, the entrant’s real answer, and all three judges’ marks.

#ModelOur score
1GPT-5.6 Sollatest100/100
2GPT-5.599/100
3Grok 4.5latest98/100
4GPT-5.3-Codex98/100
5GPT-5.6 Terra97/100
6Gemini 3.1 Pro Preview97/100
7Qwen3.7 Max95/100
8Gemini 3.5 Flash94/100
9Claude Opus 4.694/100
10DeepSeek V4 Flash93/100
11DeepSeek V4 Pro93/100
12GLM 5.293/100
13GPT-5.6 Luna93/100
14Claude Sonnet 593/100
15Mistral Medium 3.592/100
16Claude Opus 4.891/100
17Claude Fable 588/100
18Gemini 3.1 Flash Lite86/100
19Kimi K383/100

“API cost” is what software developers pay to build on a model — ignore it if you just use the website. Each model answers each task once. Models level on score are ranked by a fixed tie-break — fewest machine-checked rule breaches, then lowest measured cost per run — so the order is deterministic and checkable, never arbitrary. Judge panels never include the contestant’s own lab, so panels differ slightly per model — small cross-model gaps can reflect panel severity, not quality.

1.

GPT-5.6 Sol

100/100our pick — won on the tie-break

Made by OpenAI — their newest model. You use it inside ChatGPT — nothing to install.

Strongest showing: The wrong customer” — scored 10/10 by the panel.

The response is kind, clear, and helpful without inventing details. It perfectly addresses the situation and adheres to the under 50 words constraint (39 words).google/gemini-3.1-pro-preview, judging blind · full receipts ↓

12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $2 in / $10 out per 1M tokens · full model page →

2.

GPT-5.5

99/100

Made by OpenAI. You use it inside ChatGPT — nothing to install.

Strongest showing: The wrong customer” — scored 10/10 by the panel. Weakest: “Price rise notice” at 9/10.

The response is perfectly polite, clear, and accurately redirects the sender without inventing details. It strictly adheres to the word count constraint (41 words).google/gemini-3.1-pro-preview, judging blind · full receipts ↓

12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $5 in / $30 out per 1M tokens · full model page →

3.

Grok 4.5

98/100

Made by xAI — their newest model. You use it inside Grok — nothing to install; free via grok.com.

Strongest showing: The wrong customer” — scored 10/10 by the panel. Weakest: “The feature you won't build” at 8/10.

The response perfectly follows all instructions. It is kind, clear, correctly directs the sender without inventing details, and easily meets the strict under-50-word constraint at just 29 words.google/gemini-3.1-pro-preview, judging blind · full receipts ↓

12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $2 in / $6 out per 1M tokens · free via grok.com (checked 11 Aug 2026) · full model page →

See it for yourself

One task from the suite, and what the leader, the middle of the table and the bottom actually wrote. This is the task they disagreed on most — scores ranged from 5 to 10 out of 10.

The task — The feature you won't build

A paying customer asks when dark mode is coming; it is not on the roadmap. Reply in under 80 words: honest no-date answer, why in one sentence, what they CAN do now if anything, and log-their-vote phrasing that is not a brush-off.

Dark mode isn’t currently on our roadmap, so I can’t offer a date; we’re prioritizing other improvements based on customer impact and demand. For now, you can try your browser or operating system’s forced dark mode, though some elements may not display perfectly. I’ll log your vote against your acco…

Judge: Meets all constraints perfectly. It is 65 words, provides an honest no-date answer, explains why in one sentence, offers a practical workaround, and logs the vote professionally without sounding dismissive.

Dark mode isn't on our roadmap right now because we're prioritizing core features. In the meantime, you can use your device's system-wide dark mode or a browser extension. I've logged your vote for our team to review.

Judge: Meets the word limit and covers all requested points clearly. The reason is concise and honest. The workaround may not apply to every product context, and “for our team to review” is slightly generic, but not a clear brush-off.

> Thanks for asking — I'll be honest: dark mode isn't on our roadmap, so I can't give you a date. Our team is focused on reliability work this year, and we can only take on a few big projects at a time. Meanwhile, many customers use the Dark Reader browser extension with our web app. I've logged you…

Judge: Capped in code — the answer broke a checkable rule (113 words against a maximum of 80). The core response is excellent and meets all criteria, but the inclusion of meta-commentary pushes the total word count well over the 80-word limit constraint.

Trimmed to fit. Every full answer, every judge’s score: GPT-5.6 Sol, DeepSeek V4 Flash, Kimi K3

How this ranking is made

Every model answers the identical published task suite. Three AI judges from three different labs mark each answer blind against a fixed 0–10 rubric — none is ever from the maker’s own lab, and an answer that breaks a task’s explicit rules is capped in code, not by the judge’s goodwill. The score is the average, out of 100.

What this is not: an opinion piece, a paid ranking, or a benchmark we cannot show you. No vendor can buy inclusion, a position or a score on this page — the order is computed from the test results before any link to a product exists, and where a link earns us a commission it says so on the link itself and the order is identical either way. Every score links its raw outputs and judge verdicts. The full protocol · How we make money · receipts: GPT-5.6 Sol, GPT-5.5, Grok 4.5, GPT-5.3-Codex, GPT-5.6 Terra, Gemini 3.1 Pro Preview, Qwen3.7 Max, Gemini 3.5 Flash, Claude Opus 4.6, DeepSeek V4 Flash, DeepSeek V4 Pro, GLM 5.2, GPT-5.6 Luna, Claude Sonnet 5, Mistral Medium 3.5, Claude Opus 4.8, Claude Fable 5, Gemini 3.1 Flash Lite, Kimi K3

Questions people ask

What is the best AI for customer service replies in 2026?

GPT-5.6 Sol leads our tested ranking with 100/100 on our 12-task customer-reply suite, in a dead heat with GPT-5.5 (99). Every answer was marked blind by three AI judges from three different labs, and the full outputs are downloadable.

How is this ranking made?

Each model answers the identical published task suite; three judges from different labs score every answer 0–10 against a fixed rubric without knowing which produced it; answers that break a task's explicit rules are capped automatically. The score is the average, out of 100. No vendor pays for placement.

What happens when two models score the same?

They are separated by a fixed, published tie-break rather than by editorial choice: first the fewest machine-checked rule breaches, then the lowest measured cost per run. The order is therefore reproducible from the same data, and where the tie-break decides the top place the page says so.

How often does this page update?

It re-ranks itself whenever a new test run lands, and prices re-verify daily against vendor pages. The current ranking was last computed on 14 Aug 2026.

Head-to-heads in customer replies

Show all 20 tested pairs ▾

All comparisons →

More rankings ▾

Best AI for writing · Best AI chatbot for everyday use · Best AI for coding · Best free AI model · Best AI for spreadsheets and Excel · Best AI essay writer · Best AI for summarising documents · Best AI for extracting data from text · Best AI for writing emails · Best AI for everyday maths and percentages · Best AI for revision and study notes · Best AI for vibe coding · Best AI for making flashcards · Best AI for social media posts · Best AI for job applications and cover letters · Best AI for presentations · Best AI for writing your CV · Best AI for research skills · Best AI for creative writing · Best AI for translation · Best AI for travel planning · Best AI for emotional support · Best AI for everyday legal questions · Best AI for health questions · Best AI for writing a book · Best AI for humanising AI text · Best AI research agent · Best AI for code review · Best AI for workflow automation · Best AI for bookkeeping · Best AI for HR and employment questions · Best AI for meeting notes · Best AI for property and lettings · Best-value AI model API · every model we track · every tool