Grok 4.5 vs Qwen3.7 Max for humanising AI text

The verdict

Qwen3.7 Max wins, 9390

Same 12 humanising AI text tasks, marked blind by three rival labs. Qwen3.7 Max took 3 tasks, Grok 4.5 took 1, 8 tied. Tested 16 Aug 2026.

Where they differed most

The task — Strip the tics

Rewrite this in plain human English, keeping every fact — the date, the times, the prices: 'The village fete isn't just an event — it's a tapestry of community spirit. From the moment you arrive, you'll delve into a world of homemade cakes, vintage tractors, and the ever-popular dog show. It's not merely a fundraiser; it's a testament to what Little Barton can achieve. Saturday 12 July, 1pm till 5, entry £2, under-12s free.' One register: someone who actually lives in Little Barton telling you about it.

Here's how someone from Little Barton might actually tell you about it:

"It's our village fete on Saturday 12 July, from 1pm till 5. Entry's £2 and under-12s are free. There's homemade cakes, vintage tractors, the dog show everyone loves — the usual. It's a fundraiser, yeah, but mostly just a good chance for everyone in Little Barton to get together."
Our Little Barton village fete is coming up, and it really shows what we can achieve together as a community. It’s a fundraiser, sure, but mostly it’s just a great excuse for everyone to catch up. We’ll have homemade cakes, vintage tractors, and the dog show, which is always a massive hit. It’s on Saturday 12 July, 1pm till 5. Entry is £2, and under-12s are free. See you there!

Task by task

TaskGrok 4.5Qwen3.7 Max
The touching-base email1010
De-guru the bio1010
Product page, human voice88
Strip the tics810
The head's newsletter1010
The condolence card1010
Reply to the review, humanly810
About us, actually about us1010
Best man, not best template68
Human but still formal88
Change log required108
Refuse the disguise job1010

Full receipts: Grok 4.5, Qwen3.7 Max · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra

Questions people ask

Which is better for humanising AI text: Grok 4.5 or Qwen3.7 Max?

Qwen3.7 Max — it scored 93/100 against 90/100 on our 12-task humanising AI text suite, winning 3 tasks to 1 with 8 tied. Every answer was marked blind by three judges from three rival AI labs.

How was this tested?

Both models answered the identical published humanising AI text tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.

More humanising AI text head-to-heads: GPT-5.3-Codex vs Qwen3.7 Max · GPT-5.3-Codex vs Grok 4.5 · Claude Sonnet 5 vs Qwen3.7 Max · Claude Sonnet 5 vs Grok 4.5 · Claude Fable 5 vs Qwen3.7 Max · Claude Fable 5 vs Grok 4.5

Full ranking: Best AI for humanising AI text · model pages: Grok 4.5, Qwen3.7 Max