Best for / Extraction
Best AI for extracting data from text
If you are building anything on an AI, this is the job it does most: turn messy human text into a clean field. It is also the most objectively markable thing on the site — the JSON either parses and matches or it does not. The suite covers missing fields, ambiguous dates, unit conversion, a fixed label set, and one prompt-injection attempt to see which models follow instructions hidden in the data they were told to treat as data.
updated 13 Aug 2026 · tested by Robert Prime · re-ranks automatically when a new run lands
Scored 100/100 on our 12-task extraction suite — in a dead heat with GPT-5.3-Codex (100), so either is a fine choice.
| # | Model | Our score | Context | Free? | API $/1M in·out |
|---|---|---|---|---|---|
| 1 | DeepSeek V4 Prolatest | 100/100 | 1049k | — | $1.168 · $2.336 |
| 2 | GPT-5.3-Codex | 100/100 | 400k | — | $1.75 · $14 |
| 3 | GPT-5.6 Sollatest | 100/100 | 1050k | — | $5 · $30 |
| 4 | GPT-5.6 Terralatest | 100/100 | 1050k | — | $1 · $6 |
| 5 | Gemini 3.1 Flash Lite | 100/100 | 1049k | — | $0.25 · $1.5 |
| 6 | Gemini 3.1 Pro Preview | 100/100 | 1049k | — | $2 · $12 |
| 7 | Gemini 3.5 Flash | 100/100 | 1049k | Google AI Studio | $1.5 · $9 |
| 8 | Grok 4.5latest | 100/100 | 500k | grok.com | $2 · $6 |
| 9 | Kimi K3latest | 100/100 | 1049k | — | $3 · $15 |
| 10 | Qwen3.7 Maxlatest | 100/100 | 1000k | — | $1.475 · $4.425 |
| 11 | GPT-5.5 | 99/100 | 1050k | — | $5 · $30 |
| 12 | GPT-5.6 Lunalatest | 98/100 | 1050k | — | $0.1 · $0.6 |
| 13 | DeepSeek V4 Flashlatest | 96/100 | 1049k | — | $0.14 · $0.28 |
| 14 | GLM 5.2latest | 96/100 | 1049k | chat.z.ai | $0.5 · $1.98 |
| 15 | Claude Fable 5 | 94/100 | 1000k | — | $10 · $50 |
| 16 | Claude Opus 4.8 | 80/100 | 1000k | — | $5 · $25 |
| 17 | Mistral Medium 3.5latest | 79/100 | 262k | — | $1.5 · $7.5 |
| 18 | Claude Sonnet 5latest | 67/100 | 1000k | — | $2 · $10 |
| 19 | Claude Opus 4.6 | 59/100 | 1000k | — | $5 · $25 |
deepseek's newest model, holding about 1049k tokens of context — reached through the API.
Strongest showing: “Text to JSON” — scored 10/10 by the panel.
“Correct JSON, accurate fields, no extra prose, matches required keys exactly.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $1.168 in / $2.336 out per 1M tokens · full model page →
openai's model, holding about 400k tokens of context — reached through the API.
Strongest showing: “Text to JSON” — scored 10/10 by the panel.
“Correct JSON only, accurate extraction of all fields, matches required format exactly, concise and clear.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $1.75 in / $14 out per 1M tokens · full model page →
openai's newest model, holding about 1050k tokens of context — reached through the API.
Strongest showing: “Text to JSON” — scored 10/10 by the panel.
“Correct JSON-only output with accurate fields extracted from text, matches required keys exactly, concise and correct.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $5 in / $30 out per 1M tokens · full model page →
openai's newest model, holding about 1050k tokens of context — reached through the API.
Strongest showing: “Text to JSON” — scored 10/10 by the panel.
“Correct JSON only, accurate extraction, proper keys, no extra prose.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $1 in / $6 out per 1M tokens · full model page →
google's model, holding about 1049k tokens of context — reached through the API.
Strongest showing: “Text to JSON” — scored 10/10 by the panel.
“Correct extraction, valid JSON, only minor omission of 'Dr' title in name.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gpt-5.6-terra, grok-4.5 · API $0.25 in / $1.5 out per 1M tokens · full model page →
google's model, holding about 1049k tokens of context — reached through the API.
Strongest showing: “Text to JSON” — scored 10/10 by the panel.
“Correct JSON only, accurate extraction, proper keys, concise.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gpt-5.6-terra, grok-4.5 · API $2 in / $12 out per 1M tokens · full model page →
google's model, holding about 1049k tokens of context — reached through the API or free via Google AI Studio.
Strongest showing: “Text to JSON” — scored 10/10 by the panel.
“Correct JSON, accurate fields, clean format, no extraneous prose.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gpt-5.6-terra, grok-4.5 · API $1.5 in / $9 out per 1M tokens · free via Google AI Studio ↗ (checked 11 Aug 2026) · full model page →
x-ai's newest model, holding about 500k tokens of context — reached through the API or free via grok.com.
Strongest showing: “Text to JSON” — scored 10/10 by the panel.
“Correct JSON, accurate fields, minor stylistic issue with role phrasing ('heads clinical operations' vs 'Head of Clinical Operations'), otherwise clean and compliant.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $2 in / $6 out per 1M tokens · free via grok.com ↗ (checked 11 Aug 2026) · full model page →
moonshotai's newest model, holding about 1049k tokens of context — reached through the API.
Strongest showing: “Text to JSON” — scored 10/10 by the panel.
“Correct JSON, all fields accurate, no extra prose, exactly matches required format and content.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $3 in / $15 out per 1M tokens · full model page →
qwen's newest model, holding about 1000k tokens of context — reached through the API.
Strongest showing: “Text to JSON” — scored 10/10 by the panel.
“Correct JSON, accurate extraction, proper keys, clean formatting, no extraneous prose.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $1.475 in / $4.425 out per 1M tokens · full model page →
openai's model, holding about 1050k tokens of context — reached through the API.
Strongest showing: “Text to JSON” — scored 10/10 by the panel. Weakest: “Nested structure” at 9/10.
“Correct JSON-only output, accurate extraction of all fields, matches required format exactly, concise and clear.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $5 in / $30 out per 1M tokens · full model page →
openai's newest model, holding about 1050k tokens of context — reached through the API.
Strongest showing: “Text to JSON” — scored 10/10 by the panel. Weakest: “Classify with a fixed label set” at 8/10.
“Correct, valid JSON only, matches all facts and required keys, concise and accurate.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $0.1 in / $0.6 out per 1M tokens · full model page →
deepseek's newest model, holding about 1049k tokens of context — reached through the API.
Strongest showing: “Text to JSON” — scored 10/10 by the panel. Weakest: “Normalise inconsistent units” at 5/10. The judges flagged 2 instruction breaches across the run — those scores were capped automatically.
“Correct JSON, accurate extraction, proper keys, no extraneous prose. Minor stylistic role phrasing but accurate.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $0.14 in / $0.28 out per 1M tokens · full model page →
z-ai's newest model, holding about 1049k tokens of context — reached through the API or free via chat.z.ai.
Strongest showing: “Text to JSON” — scored 10/10 by the panel. Weakest: “Flag the ambiguity” at 5/10. The judges flagged 3 instruction breaches across the run — those scores were capped automatically.
“Correct extraction, valid JSON only, proper keys and values, no extraneous prose.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $0.5 in / $1.98 out per 1M tokens · free via chat.z.ai ↗ (checked 11 Aug 2026) · full model page →
anthropic's model, holding about 1000k tokens of context — reached through the API.
Strongest showing: “Text to JSON” — scored 10/10 by the panel. Weakest: “Flag the ambiguity” at 5/10. The judges flagged 4 instruction breaches across the run — those scores were capped automatically.
“The response perfectly follows all instructions, providing only valid JSON with the exact requested keys and accurately extracted information.”— google/gemini-3.1-pro-preview, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5 · API $10 in / $50 out per 1M tokens · full model page →
anthropic's model, holding about 1000k tokens of context — reached through the API.
Strongest showing: “Text to JSON” — scored 10/10 by the panel. Weakest: “Nested structure” at 5/10. The judges flagged 10 instruction breaches across the run — those scores were capped automatically.
“The response perfectly extracts the requested information and formats it as valid JSON with the exact specified keys, containing no additional prose.”— google/gemini-3.1-pro-preview, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5 · API $5 in / $25 out per 1M tokens · full model page →
mistralai's newest model, holding about 262k tokens of context — reached through the API.
Strongest showing: “Table from prose” — scored 10/10 by the panel. Weakest: “Nested structure” at 5/10. The judges flagged 12 instruction breaches across the run — those scores were capped automatically.
“Accurate, correctly formatted markdown table with all items, quantities, and prices correctly extracted and converted from words to numbers.”— anthropic/claude-sonnet-5, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra · API $1.5 in / $7.5 out per 1M tokens · full model page →
anthropic's newest model, holding about 1000k tokens of context — reached through the API.
Strongest showing: “Table from prose” — scored 10/10 by the panel. Weakest: “Nested structure” at 5/10. The judges flagged 19 instruction breaches across the run — those scores were capped automatically.
“Accurately extracted the items, quantities, and prices into a clean markdown table. Perfectly fulfills the task instructions.”— google/gemini-3.1-pro-preview, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5 · API $2 in / $10 out per 1M tokens · full model page →
anthropic's model, holding about 1000k tokens of context — reached through the API.
Strongest showing: “Table from prose” — scored 10/10 by the panel. Weakest: “Empty result is a valid result” at 5/10. The judges flagged 22 instruction breaches across the run — those scores were capped automatically.
“The response perfectly follows the instructions, accurately extracting the data into a correctly formatted markdown table with the specified column headers.”— google/gemini-3.1-pro-preview, judging blind · full receipts ↓
12 tasks · 13 Aug 2026 · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5 · API $5 in / $25 out per 1M tokens · full model page →
See it for yourself
One task from the suite, and what the leader, the middle of the table and the bottom actually wrote. This is the task they disagreed on most — scores ranged from 5 to 10 out of 10.
Return ONLY valid JSON, no prose, with keys name, company, role, email: 'You can reach Dr Amara Osei, who heads clinical operations at Northgate Health, on a.osei@northgate-health.co.uk.'
{
"name": "Dr Amara Osei",
"company": "Northgate Health",
"role": "Head of Clinical Operations",
"email": "a.osei@northgate-health.co.uk"
}Judge: Correct JSON, accurate fields, no extra prose, matches required keys exactly.
{
"name": "Dr Amara Osei",
"company": "Northgate Health",
"role": "Head of Clinical Operations",
"email": "a.osei@northgate-health.co.uk"
}Judge: Correct JSON, accurate extraction, proper keys, clean formatting, no extraneous prose.
```json
{
"name": "Dr Amara Osei",
"company": "Northgate Health",
"role": "heads clinical operations",
"email": "a.osei@northgate-health.co.uk"
}
```Judge: Accurately extracted all requested information into the correct JSON keys. However, the inclusion of markdown code blocks technically violates the 'ONLY valid JSON' constraint for strict parsers, though it contains no prose.
Trimmed to fit. Every full answer, every judge’s score: DeepSeek V4 Pro, Qwen3.7 Max, Claude Opus 4.6
How this ranking is made
Every model answers the identical published task suite. Three AI judges from three different labs mark each answer blind against a fixed 0–10 rubric — none is ever from the maker’s own lab, and an answer that breaks a task’s explicit rules is capped in code, not by the judge’s goodwill. The score is the average, out of 100.
What this is not: an opinion piece, an affiliate list, or a benchmark we cannot show you. Every score links its raw outputs and judge verdicts. The full protocol · receipts: DeepSeek V4 Pro, GPT-5.3-Codex, GPT-5.6 Sol, GPT-5.6 Terra, Gemini 3.1 Flash Lite, Gemini 3.1 Pro Preview, Gemini 3.5 Flash, Grok 4.5, Kimi K3, Qwen3.7 Max, GPT-5.5, GPT-5.6 Luna, DeepSeek V4 Flash, GLM 5.2, Claude Fable 5, Claude Opus 4.8, Mistral Medium 3.5, Claude Sonnet 5, Claude Opus 4.6
Questions people ask
What is the best AI for extracting data from text in 2026?
DeepSeek V4 Pro leads our tested ranking with 100/100 on our 12-task extraction suite (12 tasks), in a dead heat with GPT-5.3-Codex (100). Every answer was marked blind by three AI judges from three different labs, and the full outputs are downloadable.
How is this ranking made?
Each model answers the identical published task suite; three judges from different labs score every answer 0–10 against a fixed rubric without knowing which produced it; answers that break a task's explicit rules are capped automatically. The score is the average, out of 100. No vendor pays for placement.
How often does this page update?
It re-ranks itself whenever a new test run lands, and prices re-verify daily against vendor pages. The current ranking was last computed on 13 Aug 2026.
More rankings: Best AI for writing · Best AI chatbot for everyday use · Best AI for coding · Best free AI model · Best AI for spreadsheets and Excel · Best AI for writing essays · Best AI for summarising documents · Best-value AI model API · every model we track · every tool