Best for / Customer replies
Best AI for customer service replies
The hardest writing in business is to a customer who is angry, confused or right. This suite tests the replies that decide whether they stay: the instant refund, the clear no, the public 1-star response, the mid-outage update with no fix yet. Banned phrases are enforced — no 'unfortunately', no 'we take this seriously'. One task asks for five fake 5-star reviews; the right answer is to decline.
updated 14 Aug 2026 · tested by Robert Prime · re-ranks automatically when a new run lands
Scored 100/100 on our 12-task customer-reply suite — in a dead heat with GPT-5.5 (99), so either is a fine choice.
| # | Model | Our score | Context | Free? | API $/1M in·out |
|---|---|---|---|---|---|
| 1 | GPT-5.6 Sollatest | 100/100 | 1050k | — | $5 · $30 |
| 2 | GPT-5.5 | 99/100 | 1050k | — | $5 · $30 |
| 3 | GPT-5.3-Codex | 98/100 | 400k | — | $1.75 · $14 |
| 4 | Grok 4.5latest | 98/100 | 500k | grok.com | $2 · $6 |
| 5 | GPT-5.6 Terralatest | 97/100 | 1050k | — | $1 · $6 |
| 6 | Gemini 3.1 Pro Preview | 97/100 | 1049k | — | $2 · $12 |
| 7 | Qwen3.7 Maxlatest | 95/100 | 1000k | — | $1.475 · $4.425 |
| 8 | Claude Opus 4.6 | 94/100 | 1000k | — | $5 · $25 |
| 9 | Gemini 3.5 Flash | 94/100 | 1049k | Google AI Studio | $1.5 · $9 |
| 10 | Claude Sonnet 5latest | 93/100 | 1000k | — | $2 · $10 |
| 11 | DeepSeek V4 Flashlatest | 93/100 | 1049k | — | $0.14 · $0.28 |
| 12 | DeepSeek V4 Prolatest | 93/100 | 1049k | — | $1.168 · $2.336 |
| 13 | GLM 5.2latest | 93/100 | 1049k | chat.z.ai | $0.63 · $1.98 |
| 14 | GPT-5.6 Lunalatest | 93/100 | 1050k | — | $0.1 · $0.6 |
| 15 | Mistral Medium 3.5latest | 92/100 | 262k | — | $1.5 · $7.5 |
| 16 | Claude Fable 5 | 91/100 | 1000k | — | $10 · $50 |
| 17 | Claude Opus 4.8 | 91/100 | 1000k | — | $5 · $25 |
| 18 | Gemini 3.1 Flash Lite | 86/100 | 1049k | — | $0.25 · $1.5 |
| 19 | Kimi K3latest | 83/100 | 1049k | — | $3 · $15 |
openai's newest model, holding about 1050k tokens of context — reached through the API.
Strongest showing: “Refund the easy way” — scored 10/10 by the panel.
“Meets all constraints: under 90 words, immediate refund, timeframe stated, no forms, avoids banned phrase. Clear, concise, professional.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $5 in / $30 out per 1M tokens · full model page →
openai's model, holding about 1050k tokens of context — reached through the API.
Strongest showing: “Refund the easy way” — scored 10/10 by the panel. Weakest: “Price rise notice” at 9/10.
“Meets all constraints: under 90 words, immediate refund, timeframe given, no forms, avoids banned phrase, clear and professional.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $5 in / $30 out per 1M tokens · full model page →
openai's model, holding about 400k tokens of context — reached through the API.
Strongest showing: “Refund the easy way” — scored 10/10 by the panel. Weakest: “De-escalate a threat” at 8/10. The judges flagged 2 instruction breaches across the run — those scores were capped automatically.
“Meets all constraints: under 90 words, immediate refund, timeframe given, no forms, avoids banned phrase. Clear, polite, concise, professional-friendly tone.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $1.75 in / $14 out per 1M tokens · full model page →
x-ai's newest model, holding about 500k tokens of context — reached through the API or free via grok.com.
Strongest showing: “Refund the easy way” — scored 10/10 by the panel. Weakest: “The feature you won't build” at 8/10.
“Meets constraints: under 90 words, immediate refund, timeframe given, no forms, no banned phrase. Clear, concise, professional.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $2 in / $6 out per 1M tokens · free via grok.com ↗ (checked 11 Aug 2026) · full model page →
openai's newest model, holding about 1050k tokens of context — reached through the API.
Strongest showing: “Refund the easy way” — scored 10/10 by the panel. Weakest: “Outage update mid-incident” at 6/10.
“Meets all constraints: immediate grant, 5-7 day timeframe, no forms, no forbidden phrase, under 90 words. Clear, warm, concise, appropriate for non-technical reader.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $1 in / $6 out per 1M tokens · full model page →
google's model, holding about 1049k tokens of context — reached through the API.
Strongest showing: “Refund the easy way” — scored 10/10 by the panel. Weakest: “Apologise for your mistake” at 8/10.
“Meets all constraints: grants refund immediately, gives timeframe, no forms, no banned phrase, concise, under 90 words, polite and clear.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gpt-5.6-terra, grok-4.5 · API $2 in / $12 out per 1M tokens · full model page →
qwen's newest model, holding about 1000k tokens of context — reached through the API.
Strongest showing: “Refund the easy way” — scored 10/10 by the panel. Weakest: “Refuse the fake review request” at 8/10.
“Meets all constraints: grants refund immediately, gives timeframe, no forms, no forbidden phrase, under 90 words, polite and clear.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $1.475 in / $4.425 out per 1M tokens · full model page →
anthropic's model, holding about 1000k tokens of context — reached through the API.
Strongest showing: “Refund the easy way” — scored 10/10 by the panel. Weakest: “The wrong customer” at 5/10. The judges flagged 2 instruction breaches across the run — those scores were capped automatically.
“The response perfectly follows all instructions. It is under 90 words, grants the refund immediately, includes the 5-7 working days timeframe, mentions no forms are needed, and avoids the forbidden phrase. Highly professional and concise.”— google/gemini-3.1-pro-preview, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5 · API $5 in / $25 out per 1M tokens · full model page →
google's model, holding about 1049k tokens of context — reached through the API or free via Google AI Studio.
Strongest showing: “Refund the easy way” — scored 10/10 by the panel. Weakest: “The wrong customer” at 8/10. The judges flagged 1 instruction breach across the run — those scores were capped automatically.
“Meets constraints, concise, polite, but slightly over-explains 'no forms' which wasn't requested to mention; word count under 90, no banned phrase.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gpt-5.6-terra, grok-4.5 · API $1.5 in / $9 out per 1M tokens · free via Google AI Studio ↗ (checked 11 Aug 2026) · full model page →
anthropic's newest model, holding about 1000k tokens of context — reached through the API.
Strongest showing: “Refund the easy way” — scored 10/10 by the panel. Weakest: “Say no without weasel words” at 5/10. The judges flagged 4 instruction breaches across the run — those scores were capped automatically.
“The response perfectly follows all instructions. It is under 90 words, grants the refund immediately, includes the correct timeframe, avoids forms, and omits the forbidden phrase. The tone is highly professional and empathetic.”— google/gemini-3.1-pro-preview, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5 · API $2 in / $10 out per 1M tokens · full model page →
deepseek's newest model, holding about 1049k tokens of context — reached through the API.
Strongest showing: “Refund the easy way” — scored 10/10 by the panel. Weakest: “Refuse the fake review request” at 8/10.
“Meets all constraints: under 90 words, immediate refund, timeframe given, no forms, no forbidden phrase. Clear, polite, concise, professional-friendly.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $0.14 in / $0.28 out per 1M tokens · full model page →
deepseek's newest model, holding about 1049k tokens of context — reached through the API.
Strongest showing: “Refund the easy way” — scored 10/10 by the panel. Weakest: “The wrong customer” at 8/10. The judges flagged 2 instruction breaches across the run — those scores were capped automatically.
“Meets all constraints: under 90 words, immediate refund, timeframe given, no forms, avoids banned phrase, professional and clear.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $1.168 in / $2.336 out per 1M tokens · full model page →
z-ai's newest model, holding about 1049k tokens of context — reached through the API or free via chat.z.ai.
Strongest showing: “Refund the easy way” — scored 10/10 by the panel. Weakest: “Refuse the fake review request” at 7/10.
“Meets constraints: under 90 words, immediate grant, timeframe given, no forms, avoids banned phrase. Clear, polite, concise; minor stylistic redundancy only.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $0.63 in / $1.98 out per 1M tokens · free via chat.z.ai ↗ (checked 11 Aug 2026) · full model page →
openai's newest model, holding about 1050k tokens of context — reached through the API.
Strongest showing: “Refund the easy way” — scored 10/10 by the panel. Weakest: “Outage update mid-incident” at 5/10.
“Meets all constraints: under 90 words, immediate grant, timeframe given, no forms, no banned phrase. Clear, polite, concise.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $0.1 in / $0.6 out per 1M tokens · full model page →
mistralai's newest model, holding about 262k tokens of context — reached through the API.
Strongest showing: “Refund the easy way” — scored 10/10 by the panel. Weakest: “The feature you won't build” at 6/10. The judges flagged 2 instruction breaches across the run — those scores were capped automatically.
“Concise, grants refund immediately, gives timeframe, avoids forbidden phrase and forms, under 90 words. Minor: doesn't explicitly mention photos, but not required.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $1.5 in / $7.5 out per 1M tokens · full model page →
anthropic's model, holding about 1000k tokens of context — reached through the API.
Strongest showing: “Refund the easy way” — scored 10/10 by the panel. Weakest: “The confused customer” at 5/10. The judges flagged 6 instruction breaches across the run — those scores were capped automatically.
“The response perfectly follows all instructions and constraints. It is under 90 words, grants the refund immediately, includes the timeframe, avoids forms and the banned phrase, and is highly professional and clear.”— google/gemini-3.1-pro-preview, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5 · API $10 in / $50 out per 1M tokens · full model page →
anthropic's model, holding about 1000k tokens of context — reached through the API.
Strongest showing: “Refund the easy way” — scored 10/10 by the panel. Weakest: “The wrong customer” at 5/10. The judges flagged 6 instruction breaches across the run — those scores were capped automatically.
“The response perfectly follows all instructions. It is empathetic, clear, grants the refund immediately, includes the timeframe, and strictly adheres to the word count and negative constraints.”— google/gemini-3.1-pro-preview, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5 · API $5 in / $25 out per 1M tokens · full model page →
google's model, holding about 1049k tokens of context — reached through the API.
Strongest showing: “Refund the easy way” — scored 10/10 by the panel. Weakest: “Chargeback rebuttal facts” at 6/10. The judges flagged 1 instruction breach across the run — those scores were capped automatically.
“Meets all constraints: grants refund immediately, gives 5-7 day window, no forms, no banned phrase, under 90 words, clear and professional.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gpt-5.6-terra, grok-4.5 · API $0.25 in / $1.5 out per 1M tokens · full model page →
moonshotai's newest model, holding about 1049k tokens of context — reached through the API.
Strongest showing: “Refund the easy way” — scored 10/10 by the panel. Weakest: “The wrong customer” at 5/10. The judges flagged 6 instruction breaches across the run — those scores were capped automatically.
“Meets all constraints: under 90 words (~84), immediate refund granted, timeframe given, no forms, no banned phrase. Clear, warm, concise, professional.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 14 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $3 in / $15 out per 1M tokens · full model page →
See it for yourself
One task from the suite, and what the leader, the middle of the table and the bottom actually wrote. This is the task they disagreed on most — scores ranged from 5 to 10 out of 10.
A customer wants a refund on a personalised item after 60 days; policy is 30 days and the item was made to order. Write the reply: under 110 words, the no must be clear in the first two sentences, offer one genuine alternative, never say 'unfortunately' or 'we regret'.
We can’t issue a refund because the request was made after our 30-day return period. Personalised, made-to-order items are also non-refundable. If the item is faulty, please send us photos and a description of the issue, and we can arrange a repair. Thank you for your understanding.
Judge: Clear no in first two sentences, offers repair alternative, avoids banned words, concise and under word limit. Minor: alternative slightly conditional but genuine.
Subject: Your Refund Request Hi [Name], Thank you for reaching out. I'm not able to approve a refund for this item—as a personalised, made-to-order piece, it falls outside our 30-day policy, and 60 days have now passed. I do want to help find a solution. If there's a specific issue with the item…
Judge: The response is professional, clear, and follows the negative constraints and alternative requirement. However, it violates the length constraint, coming in at approximately 116 words (over the 110-word limit).
Hi [Customer name], Thanks for getting in touch. We're not able to offer a refund on this order — it's a personalised, made-to-order item, and it's now outside our 30-day returns window, so it isn't eligible under our policy. We'd still like to do something for you. We can offer 20% off your next…
Judge: Clear no in first two sentences, offers genuine alternative, avoids banned words, under 110 words, professional tone. Minor stylistic tweaks possible but solid.
Trimmed to fit. Every full answer, every judge’s score: GPT-5.6 Sol, Claude Sonnet 5, Kimi K3
How this ranking is made
Every model answers the identical published task suite. Three AI judges from three different labs mark each answer blind against a fixed 0–10 rubric — none is ever from the maker’s own lab, and an answer that breaks a task’s explicit rules is capped in code, not by the judge’s goodwill. The score is the average, out of 100.
What this is not: an opinion piece, an affiliate list, or a benchmark we cannot show you. Every score links its raw outputs and judge verdicts. The full protocol · receipts: GPT-5.6 Sol, GPT-5.5, GPT-5.3-Codex, Grok 4.5, GPT-5.6 Terra, Gemini 3.1 Pro Preview, Qwen3.7 Max, Claude Opus 4.6, Gemini 3.5 Flash, Claude Sonnet 5, DeepSeek V4 Flash, DeepSeek V4 Pro, GLM 5.2, GPT-5.6 Luna, Mistral Medium 3.5, Claude Fable 5, Claude Opus 4.8, Gemini 3.1 Flash Lite, Kimi K3
Questions people ask
What is the best AI for customer service replies in 2026?
GPT-5.6 Sol leads our tested ranking with 100/100 on our 12-task customer-reply suite (12 tasks), in a dead heat with GPT-5.5 (99). Every answer was marked blind by three AI judges from three different labs, and the full outputs are downloadable.
How is this ranking made?
Each model answers the identical published task suite; three judges from different labs score every answer 0–10 against a fixed rubric without knowing which produced it; answers that break a task's explicit rules are capped automatically. The score is the average, out of 100. No vendor pays for placement.
How often does this page update?
It re-ranks itself whenever a new test run lands, and prices re-verify daily against vendor pages. The current ranking was last computed on 14 Aug 2026.
More rankings: Best AI for writing · Best AI chatbot for everyday use · Best AI for coding · Best free AI model · Best AI for spreadsheets and Excel · Best AI for writing essays · Best AI for summarising documents · Best AI for extracting data from text · Best AI for writing emails · Best AI for everyday maths and percentages · Best AI for revision and study notes · Best-value AI model API · every model we track · every tool