Best for / Best-value API

Best-value AI model API

For anyone building on AI, the question is capability per pound. We score each model on the same 30-task general suite and put the score beside its live per-million-token price — which for some models moves during the day, because OpenRouter blends prices across providers. Where a price is volatile we say so.

updated 11 Aug 2026 · tested by Robert Prime · re-ranks automatically when a new run lands

The verdict

GLM 5.2

Scored 92/100 on our best-value api testsin a dead heat with GPT-5.6 Terra (92), so either is a fine choice. Free to use via chat.z.ai.

#ModelOur scoreSentimentFree?API $/1M in·out
1GLM 5.2latest92/100chat.z.ai$0.76 · $1.5356
2GPT-5.6 Terralatest92/100$1 · $6
3Grok 4.5latest92/100grok.com$2 · $6
4Kimi K3latest91/100$3 · $15
5DeepSeek V4 Prolatest88/100$0.6317 · $1.2634
6Gemini 3.1 Pro Preview88/100$2 · $12
7Claude Opus 4.685/100$5 · $25
1.

GLM 5.2

92/100our pick

z-ai's newest model, holding about 1049k tokens of context — reached through the API or free via chat.z.ai.

Strongest showing: Cold email” — scored 10/10 by the panel. Weakest: “Long generation” at 5/10. The judges flagged 4 instruction breaches across the run — those scores were capped automatically.

Clear, friendly, professional, single CTA, no buzzwords, under 120 words (~99). Minor genericness but well-structured and appropriate for target audience.anthropic/claude-sonnet-5, judging blind · full receipts ↓

30 tasks · 11 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $0.76 in / $1.5356 out per 1M tokens · free via chat.z.ai (checked 11 Aug 2026) · full model page →

openai's newest model, holding about 1050k tokens of context — reached through the API.

Strongest showing: Summarise messy notes” — scored 10/10 by the panel. Weakest: “Long generation” at 5/10. The judges flagged 9 instruction breaches across the run — those scores were capped automatically.

Accurate, concise, exactly 5 bullets, captures all key points, readable in 20 seconds, director-friendly formatting.anthropic/claude-sonnet-5, judging blind · full receipts ↓

30 tasks · 11 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $1 in / $6 out per 1M tokens · full model page →

3.

Grok 4.5

92/100

x-ai's newest model, holding about 500k tokens of context — reached through the API or free via grok.com.

Strongest showing: Cold email” — scored 10/10 by the panel. Weakest: “Long generation” at 5/10. The judges flagged 5 instruction breaches across the run — those scores were capped automatically.

Meets word limit, friendly tone, clear CTA, no buzzwords; minor placeholder-heavy structure but otherwise strong and professional.anthropic/claude-sonnet-5, judging blind · full receipts ↓

30 tasks · 11 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $2 in / $6 out per 1M tokens · free via grok.com (checked 11 Aug 2026) · full model page →

4.

Kimi K3

91/100

moonshotai's newest model, holding about 1049k tokens of context — reached through the API.

Strongest showing: Summarise messy notes” — scored 10/10 by the panel. Weakest: “Long generation” at 5/10. The judges flagged 7 instruction breaches across the run — those scores were capped automatically.

Accurate, concise, exactly 5 bullets, captures all key points, easily readable in 20 seconds.anthropic/claude-sonnet-5, judging blind · full receipts ↓

30 tasks · 11 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $3 in / $15 out per 1M tokens · full model page →

deepseek's newest model, holding about 1049k tokens of context — reached through the API.

Strongest showing: Cold email” — scored 10/10 by the panel. Weakest: “Long generation” at 5/10. The judges flagged 9 instruction breaches across the run — those scores were capped automatically.

Meets word limit, clear CTA, friendly/professional tone, no buzzwords. Minor: could be slightly more brewery-specific, but overall strong and concise.anthropic/claude-sonnet-5, judging blind · full receipts ↓

30 tasks · 11 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $0.6317 in / $1.2634 out per 1M tokens · full model page →

google's model, holding about 1049k tokens of context — reached through the API.

Strongest showing: Maths word problem” — scored 10/10 by the panel. Weakest: “Medium generation” at 3/10. The judges flagged 9 instruction breaches across the run — those scores were capped automatically.

Correct calculations, clear step-by-step working, matches expected answer of 10:25, well-formatted and easy to follow.anthropic/claude-sonnet-5, judging blind · full receipts ↓

30 tasks · 11 Aug 2026 · judges claude-sonnet-5, gpt-5.6-terra, grok-4.5 · API $2 in / $12 out per 1M tokens · full model page →

anthropic's model, holding about 1000k tokens of context — reached through the API.

Strongest showing: Cold email” — scored 10/10 by the panel. Weakest: “Long generation” at 5/10. The judges flagged 20 instruction breaches across the run — those scores were capped automatically.

The response perfectly follows all instructions. It is under the 120-word limit (approx 107 words), includes the required details, maintains a friendly yet professional tone, avoids buzzwords, and ends with a single clear call to action.google/gemini-3.1-pro-preview, judging blind · full receipts ↓

30 tasks · 11 Aug 2026 · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5 · API $5 in / $25 out per 1M tokens · full model page →

How this ranking is made

Every model answers the identical published task suite. Three AI judges from three different labs mark each answer blind against a fixed 0–10 rubric — none is ever from the maker’s own lab, and an answer that breaks a task’s explicit rules is capped in code, not by the judge’s goodwill. The score is the average, out of 100.

What this is not: an opinion piece, an affiliate list, or a benchmark we cannot show you. Every score links its raw outputs and judge verdicts. The full protocol · receipts: GLM 5.2, GPT-5.6 Terra, Grok 4.5, Kimi K3, DeepSeek V4 Pro, Gemini 3.1 Pro Preview, Claude Opus 4.6

Questions people ask

What is the best-value ai model api in 2026?

GLM 5.2 leads our tested ranking with 92/100 across 30 best-value api tasks, in a dead heat with GPT-5.6 Terra (92). Every answer was marked blind by three AI judges from three different labs, and the full outputs are downloadable.

How is this ranking made?

Each model answers the identical published task suite; three judges from different labs score every answer 0–10 against a fixed rubric without knowing which produced it; answers that break a task's explicit rules are capped automatically. The score is the average, out of 100. No vendor pays for placement.

How often does this page update?

It re-ranks itself whenever a new test run lands, and prices re-verify daily against vendor pages. The current ranking was last computed on 11 Aug 2026.

Is GLM 5.2 free?

Yes — GLM 5.2 can be used at no cost via chat.z.ai, confirmed on 11 Aug 2026.

Head-to-head in best-value api

More rankings: Best AI for writing · Best AI chatbot for everyday use · Best AI for coding · Best free AI model · every model we track · every tool