Best AI for humanising AI text

The verdict

GPT-5.3-Codex

Scored 97/100 on our 12-task humanising suitea single point ahead of Claude Sonnet 5 (96) — effectively level.

What to actually do

GPT-5.3-Codex is the engine inside ChatGPT. Go to chatgpt.comthe free tier is fine to start. Paid plans start at £7/month (go plan, vendor’s own price). The free tier is ChatGPT’s, not a promise about this exact model — we haven’t verified which plan carries it. Not fussed about the last point or two? Any of the top 2 here will serve you well.

Two things are marked and both must hold: does the rewrite sound like a person, and did every fact survive. The suite feeds it the classics — 'I hope this email finds you well', the visionary-thought-leader bio, the tapestry-of-community-spirit newsletter — plus the trap of over-correcting a formal planning objection into chattiness. No detector-beating claims anywhere: one task asks to disguise coursework so 'the checker can't tell', and the right answer is to refuse the disguise and offer legitimate help instead.

updated 16 Aug 2026 · tested by Robert Prime · re-ranks automatically when a new run lands

#ModelOur score
1GPT-5.3-Codex97/100
2Claude Sonnet 5latest96/100
3Claude Fable 595/100
4GPT-5.6 Sol94/100
5GPT-5.594/100
6GLM 5.293/100
7GPT-5.6 Terra93/100
8Qwen3.7 Max93/100
9GPT-5.6 Luna92/100
10Grok 4.590/100
11Claude Opus 4.890/100
12DeepSeek V4 Pro89/100
13Claude Opus 4.689/100
14DeepSeek V4 Flash88/100
15Gemini 3.5 Flash87/100
16Gemini 3.1 Pro Preview87/100
17Kimi K385/100
18Mistral Medium 3.581/100
19Gemini 3.1 Flash Lite80/100

“API cost” is what software developers pay to build on a model — ignore it if you just use the website. Each model answers each task once. Models level on score are ranked by a fixed tie-break — fewest machine-checked rule breaches, then lowest measured cost per run — so the order is deterministic and checkable, never arbitrary. Judge panels never include the contestant’s own lab, so panels differ slightly per model — small cross-model gaps can reflect panel severity, not quality.

1.

GPT-5.3-Codex

97/100our pick — won on the tie-break

Made by OpenAI. You use it inside ChatGPT — nothing to install.

Strongest showing: Product page, human voice” — scored 10/10 by the panel.

The response perfectly captures the requested market stall tone while naturally incorporating every required specification. It is concise, clear, and highly effective for the target audience.google/gemini-3.1-pro-preview, judging blind · full receipts ↓

12 tasks · 16 Aug 2026 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5 · API $1.75 in / $14 out per 1M tokens · full model page →

Made by Anthropic — their newest model. You use it inside Claude — nothing to install.

Strongest showing: Product page, human voice” — scored 10/10 by the panel. Weakest: “The head's newsletter” at 9/10. On one task a judge flagged a possible instruction breach — advisory only: the machine check did not confirm it and the panel outvoted it.

The response perfectly captures the requested market stall tone while accurately including every required specification. It is concise, natural, and highly effective.google/gemini-3.1-pro-preview, judging blind · full receipts ↓

12 tasks · 16 Aug 2026 · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5 · API $2 in / $10 out per 1M tokens · full model page →

Made by Anthropic. You use it inside Claude — nothing to install.

Strongest showing: The head's newsletter” — scored 10/10 by the panel. Weakest: “The touching-base email” at 9/10. On one task a judge flagged a possible instruction breach — advisory only: the machine check did not confirm it and the panel outvoted it.

The response perfectly adopts a warm, direct head teacher's voice while retaining all required facts (old date, new date, departure time, and wellies). It is clear, concise, and highly useful.google/gemini-3.1-pro-preview, judging blind · full receipts ↓

12 tasks · 16 Aug 2026 · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5 · API $10 in / $50 out per 1M tokens · full model page →

See it for yourself

One task from the suite, and what the leader, the middle of the table and the bottom actually wrote. This is the task they disagreed on most — scores ranged from 6 to 10 out of 10.

The task — The head's newsletter

Rewrite for the school newsletter in a head teacher's human voice — warm, direct, every fact kept: the old date, the new date, the departure time, the wellies: 'Please be advised that due to unforeseen circumstances relating to ground conditions, the Year 4 farm visit has been rescheduled from Tuesday 8 October to Thursday 17 October. Departure will occur at 9.15am promptly. It is imperative that all pupils are equipped with wellington boots on the day in question. We apologise for any inconvenience caused and thank you for your continued understanding and support.'

Dear families,

A quick update about the Year 4 farm visit: because of unforeseen ground conditions, we’ve had to move the trip from **Tuesday 8 October** to **Thursday 17 October**.  
Please make sure children are ready to leave **promptly at 9.15am** on the day, and that every pupil comes in **wel…

Judge: The response perfectly adopts a warm, direct tone suitable for a head teacher while retaining all required facts accurately. It is clear, concise, and highly useful.

**Year 4 Farm Visit – Date Change**

Dear parents and carers,

Just a quick note about the Year 4 farm visit. Because of the ground conditions, we’ve had to move it from Tuesday 8 October to Thursday 17 October. 

We’ll still be leaving promptly at 9.15am, so please make sure everyone is here on tim…

Judge: The response perfectly captures a warm, direct head teacher's voice while retaining all required facts (old date, new date, departure time, and wellies). It is clear, concise, and highly useful.

Here are a few options, depending on the tone of your newsletter:

**Option 1: Warm and friendly (Best for a general newsletter)**
"I’m writing to let you know that we’ve had to move our Year 4 farm visit. Due to the recent weather, the ground at the farm is currently a bit too muddy for us to visit…

Judge: Facts preserved in warm voice, but multiple options instead of one rewrite; added unstated weather details.

Trimmed to fit. Every full answer, every judge’s score: GPT-5.3-Codex, Grok 4.5, Gemini 3.1 Flash Lite

How this ranking is made

Every model answers the identical published task suite. Three AI judges from three different labs mark each answer blind against a fixed 0–10 rubric — none is ever from the maker’s own lab, and an answer that breaks a task’s explicit rules is capped in code, not by the judge’s goodwill. The score is the average, out of 100.

What this is not: an opinion piece, an affiliate list, or a benchmark we cannot show you. Every score links its raw outputs and judge verdicts. The full protocol · receipts: GPT-5.3-Codex, Claude Sonnet 5, Claude Fable 5, GPT-5.6 Sol, GPT-5.5, GLM 5.2, GPT-5.6 Terra, Qwen3.7 Max, GPT-5.6 Luna, Grok 4.5, Claude Opus 4.8, DeepSeek V4 Pro, Claude Opus 4.6, DeepSeek V4 Flash, Gemini 3.5 Flash, Gemini 3.1 Pro Preview, Kimi K3, Mistral Medium 3.5, Gemini 3.1 Flash Lite

Questions people ask

What is the best AI for humanising AI text in 2026?

GPT-5.3-Codex leads our tested ranking with 97/100 on our 12-task humanising suite (12 tasks), in a dead heat with Claude Sonnet 5 (96). Every answer was marked blind by three AI judges from three different labs, and the full outputs are downloadable.

How is this ranking made?

Each model answers the identical published task suite; three judges from different labs score every answer 0–10 against a fixed rubric without knowing which produced it; answers that break a task's explicit rules are capped automatically. The score is the average, out of 100. No vendor pays for placement.

How often does this page update?

It re-ranks itself whenever a new test run lands, and prices re-verify daily against vendor pages. The current ranking was last computed on 16 Aug 2026.

Head-to-heads in humanising ai text

Show all 20 tested pairs ▾

All comparisons →

More rankings ▾

Best AI for writing · Best AI chatbot for everyday use · Best AI for coding · Best free AI model · Best AI for spreadsheets and Excel · Best AI essay writer · Best AI for summarising documents · Best AI for extracting data from text · Best AI for writing emails · Best AI for everyday maths and percentages · Best AI for customer service replies · Best AI for revision and study notes · Best AI for vibe coding · Best AI for making flashcards · Best AI for social media posts · Best AI for job applications and cover letters · Best AI for presentations · Best AI for writing your CV · Best AI for research skills · Best AI for creative writing · Best AI for translation · Best AI for travel planning · Best AI for emotional support · Best AI for everyday legal questions · Best AI for health questions · Best AI for writing a book · Best-value AI model API · every model we track · every tool