Battles / Best-value API

Probably GLM 5.2

GLM 5.2 vs Claude Opus 4.6 · Best-value API

GLM 5.2 came out in front, but not by enough for us to call it proven on a suite this size. Treat it as the way to bet, not as a settled result.

83
of 11 decided · 19 tied

GLM 5.2 took 8 of the 11 tasks that had a clear winner (GLM 5.2 8, Claude Opus 4.6 3). The judge could pick a winner on 11 of 30 tasks; on the other 19 it could not tell them apart. That is a lean, not a proven win — at this sample size we cannot rule out chance, so we are not calling it decisive.

So which should you pick?

Worth knowing: Claude Opus 4.6 answered faster, GLM 5.2 costs less.

Claude Opus 4.6 answered faster - 5.0s against 11.0sGLM 5.2 was 1.6x cheaper to run - $0.1314 against $0.2052 across the whole suite. Treat that with caution: GLM 5.2's API price moved 329% in the last 7 days ($0.2772-$1.19 per million input tokens), because the provider blends it across whoever is serving the model
tested 7 Aug 202630 tasks3 judges from different labs
Show the full workings
last verified 7 Aug 2026previously GLM 5.2 91 Claude Opus 4.6, published 7 Aug 2026 under judge-2026-08 — replaced: re-run under judge-2026-08suite suite-2026-07judges: google/gemini-3.1-pro-preview, openai/gpt-5.6-terra, x-ai/grok-4.5 (none of them a contestant)judge protocol judge-2026-08 — constraint cap enforced in code, applied to 71 judge scores in this runnot statistically decisive — Wilcoxon signed-rank on score margins p=0.0198 (n=23); sign test on win counts p=0.2266 — held to our confidence gatejudge swap-consistency (Cohen’s κ) = 0.55moderate agreement κ measures how often a judge reached the same verdict when the two answers were shown in the opposite order, corrected for agreement you would get by chance. A battle below 0.21 does not publish a winner at all; above it, this number is how much of the scoreline survived the swap test rather than a claim that the judges were reliable.

The evidence

Suite-by-suite
Writing
140
Coding
131
Reasoning
041
Extraction
230
Instruction-following
320
Speed & cost
131

blue = GLM 5.2 wins · grey = ties · white = Claude Opus 4.6 wins (5 tasks per suite)

Round-by-round — all 30 tasks
tieCold emailWriting · 8.8 v 9.3Both emails follow all instructions and constraints perfectly. Response 2 wins because its tone is slightly more natural and engaging…
PROMPT

Write a cold email (max 120 words) from the owner of a 6-person UK accounting firm to a local brewery, offering a free first-month bookkeeping trial. Friendly but professional, one clear call to action, no buzzwords.

GLM 5.2 · 9.6s · $0.0029

Subject: Free first month of bookkeeping for [Brewery Name] Hi [Name], I’m [Your Name], owner of [Firm Name], a six-person accounting firm based locally. We help businesses like [Brewery Name] keep their finances tidy so you can focus on brewing. Tracking inventory, daily expenses, and cash flow can be a lot to handle alongside running a brewery. To show you how we can help, I’d like to offer a free first month of our bookkeeping service, with no obligation to continue. Would you be open to a brief 10-minute call next week to see if we’re a good fit? Cheers, [Your Name] [Firm Name]

Claude Opus 4.6 · 6.3s · $0.0047

**Subject: Happy to help with your books – first month on us** Hi, I'm Sarah, owner of Clearpath Accounting here in [Town]. We're a small team of six, and we work with several food and drink businesses locally. I know running a brewery means your time goes into making great beer, not reconciling invoices. We'd like to offer you a free month of bookkeeping – no strings, no contract. It's simply how we prefer to introduce ourselves. If the fit works, we can talk about continuing. If not, you walk away with a tidy set of books for the month. Would you be open to a 15-minute chat this week or next? Best, Sarah Mitchell Clearpath Accounting [phone] | [email]

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: B — Both emails follow all instructions and constraints perfectly. Response 2 wins because its tone is slightly more natural and engaging, specifically the line about 'making great beer, not reconciling invoices,' which perfectly captures the friendly but professional vibe requested. / Both responses follow all instructions and constraints, staying under the 120-word limit. Response 1 is slightly better written, feeling more natural and less reliant on placeholders, making it more useful as a ready-to-send template. | openai/gpt-5.6-terra: tie (flipped on swap) — Both meet the 120-word limit, tone, offer, and single-call-to-action requirements. Response 2 is more tailored to a brewery, clearly explains the low-risk trial, and sounds more natural and persuasive. Response 1 is strong but slightly more generic. / Both meet the 120-word limit, use a friendly professional tone, offer a free first month, and include one clear call to action. Response 2 is slightly more tailored to brewery bookkeeping needs and avoids the potentially generic claim about existing food-and-drink clients. | x-ai/grok-4.5: tie (flipped on swap) — Both meet max 120 words, 6-person firm, free trial offer, friendly-pro tone, single CTA, no buzzwords. Equally clear, useful, accurate. / Both meet max 120 words, offer free trial, one CTA, friendly-pro tone, no buzzwords, mention 6-person firm. R1 more natural, brewery-specific, polished; R2 solid but more generic placeholders.

GLM 5.2Product descriptionWriting · 9.5 v 8.5Both responses meet all constraints, including the 60-80 word count and avoiding the cliché. Response 1 flows slightly better and feels more…
PROMPT

Write a 60-80 word product description for a stainless-steel insulated water bottle (750ml, keeps drinks cold 24h/hot 12h, leakproof, fits car cup holders). Target audience: gym-goers. Avoid cliches like 'stay hydrated in style'.

GLM 5.2 · 17.3s · $0.0052

Fuel your training sessions with our 750ml stainless-steel water bottle. Engineered for the gym, its double-wall insulation keeps your pre-workout ice-cold for 24 hours and your post-gym coffee hot for 12. The leakproof lid lets you toss it straight into your gym bag without worry. Designed to fit car cup holders, it easily transitions from your commute to the weight room floor. Durable, reliable, and ready for heavy lifts.

Claude Opus 4.6 · 5.4s · $0.0033

# HydroCore 750ml Insulated Water Bottle Engineered for athletes who demand more from their gear. This double-wall stainless-steel bottle locks in ice-cold refreshment for 24 hours or keeps your pre-workout hot for 12. The leakproof seal means it tosses into any gym bag without worry, while the slim profile slides right into car cup holders between sessions. Durable, sweat-free, and built to survive every rep, set, and sprint you throw at it.

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: A — Both responses meet all constraints, including the 60-80 word count and avoiding the cliché. Response 1 flows slightly better and feels more natural as a product description, seamlessly integrating the features with the gym-goer lifestyle. / Both responses follow all instructions and constraints perfectly, including the 60-80 word limit. Response 2 flows slightly better as a cohesive product description without relying on a title, making it highly effective for the target audience. | openai/gpt-5.6-terra: A — Response 1 is accurate, clearly gym-focused, and includes every required feature in 69 words without clichés. Response 2 also meets length and feature requirements, but adds unsupported “sweat-free” and extreme durability claims. Its heading is unnecessary, and a few phrases are less precise. / Both meet the 60–80 word requirement, include all specified features, target gym-goers, and avoid the banned cliché. Response 2 is slightly clearer and more polished, with a stronger practical progression from commute to workout. Response 1 is effective but includes less-supported claims such as “sweat-free” and “built to survive.” | x-ai/grok-4.5: A — Both hit 60-80 words, cover all specs, target gym users, avoid banned cliche. R1 is tighter pure description; R2 adds unneeded title/brand. / Both hit 60-80 words, cover all specs, target gym-goers, avoid banned cliche. R2 slightly tighter and more natural flow for busy pros.

tieSummarise messy notesWriting · 9.3 v 9.2Both responses accurately summarize the notes into exactly 5 bullets. Response 1 is slightly more concise and scannable, making it better suited…
PROMPT

Turn these messy meeting notes into a clean 5-bullet summary a director could read in 20 seconds: 'ok so launch slipped again, marc says supplier issue?? maybe 3 wks. sarah wants to push the email campaign anyway but tom thinks wait. budget - we're 8k over, mostly the packaging redo. Q: do we tell retail partners now or after new date confirmed. also NEED to hire the warehouse temp before august rush. next mtg tues.'

GLM 5.2 · 5.7s · $0.0030

* **Launch:** Delayed ~3 weeks due to supplier issues. * **Marketing:** Email campaign timing disputed (launch now vs. wait for new date). * **Budget:** $8k over budget, primarily driven by packaging redesign. * **Decision Needed:** When to notify retail partners about the delay. * **Action Items:** Hire warehouse temp before August; next meeting Tuesday.

Claude Opus 4.6 · 5.0s · $0.0037

• **Launch delayed ~3 weeks** due to a supplier issue (details to be confirmed by Marc). • **Email campaign timing undecided** — Sarah recommends proceeding now; Tom advises holding until the new launch date is set. • **Budget is $8K over**, driven primarily by the packaging redesign. • **Retail partner communication pending** — team needs a decision on whether to notify them now or wait until the revised date is locked in. • **Warehouse temp hire is urgent** — must be in place before the August rush. Next meeting Tuesday.

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: A — Both responses accurately summarize the notes into exactly 5 bullets. Response 1 is slightly more concise and scannable, making it better suited for a busy director to read in 20 seconds. / Both responses accurately summarize the notes into exactly 5 bullets. Response 2 is slightly better for a busy director because the bolded category headers make it exceptionally scannable and concise. | openai/gpt-5.6-terra: tie (flipped on swap) — Both provide exactly five clear, director-ready bullets and preserve the key issues. Response 2 is stronger because it attributes the differing marketing views, clearly frames the retail communication decision, and retains urgency and ownership context without adding meaningful clutter. / Both accurately produce five director-ready bullets and preserve all key issues, decisions, actions, and timing. Response 2 is slightly more concise and scannable for a 20-second read while retaining the necessary detail. Response 1 is also strong but marginally more verbose. | x-ai/grok-4.5: tie (flipped on swap) — Both deliver accurate, clean 5-bullet summaries matching notes. Equally follow format, useful and clear for a director; minor style differences only. / Both accurately cover all key points in exactly 5 clean bullets. Resp1 retains slightly more nuance and decision context useful to a director; Resp2 is tighter but loses minor detail. Equal on format compliance.

tieTone rewriteWriting · 9.7 v 9.3Both responses successfully rewrite the complaint reply to be warm and take responsibility while retaining all original facts. Both are well under…
PROMPT

Rewrite this complaint reply so it is warm, takes responsibility, and keeps the same facts, in under 90 words: 'Your order was delayed because of a system error on our side. The refund has been processed and will arrive in 3-5 days. We cannot offer further compensation. Let us know if you have questions.'

GLM 5.2 · 9.8s · $0.0026

I am so sorry for the delay with your order. We experienced a system error on our end, and we take full responsibility for this inconvenience. Your refund has been processed and will arrive in 3-5 days. While we are unable to offer further compensation at this time, we truly appreciate your patience and understanding. Please let us know if you have any questions—we are always here to help!

Claude Opus 4.6 · 3.7s · $0.0030

We're truly sorry for the delay with your order — this was caused by a system error on our end, and we take full responsibility for the inconvenience. We've processed your refund, which should arrive within 3–5 days. While we're unable to offer additional compensation at this time, please know we value your patience and your business. Don't hesitate to reach out if you have any questions — we're here to help and want to make sure you're taken care of.

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: tie (flipped on swap) — Both responses successfully rewrite the complaint reply to be warm and take responsibility while retaining all original facts. Both are well under the 90-word limit and are highly useful and clear. / Both responses follow all instructions perfectly, keeping the facts while adding warmth and taking responsibility. Response 2 is slightly better structured with a paragraph break, making it more readable for a busy professional. | openai/gpt-5.6-terra: tie (flipped on swap) — Both responses preserve all facts, clearly accept responsibility, use a warm professional tone, and remain well under 90 words. Both are concise and easy for a customer to understand. Response 2 is slightly more personalized, while Response 1 is slightly more direct; neither has a material advantage. / Both are under 90 words, warm, accountable, and retain the required facts. Response 2 is slightly more concise and closer to the original wording. Response 1 is also strong but adds broader assurances about being “taken care of,” which is slightly less precise. | x-ai/grok-4.5: tie — Both warm, take responsibility, preserve all facts exactly, under 90 words, clear and useful. Nearly identical quality. / Both rewrites are warm, accept responsibility, preserve all facts exactly, and stay well under 90 words. Equally clear, concise, and useful.

tieHeadlinesWriting · 9.3 v 8.8Both responses followed all instructions and constraints perfectly. Response 2 wins because its curiosity-driven headlines are much more engaging…
PROMPT

Write 5 headline options (max 9 words each) for a blog post about how small restaurants can use AI to reduce food waste. Mix: 2 practical, 2 curiosity-driven, 1 with a number.

GLM 5.2 · 11.0s · $0.0033

Here are 5 headline options for your blog post, all 9 words or less: **Practical** 1. How Small Restaurants Can Use AI to Cut Waste 2. Slash Food Waste in Small Restaurants Using Simple AI **Curiosity-Driven** 3. The AI Secret Saving Small Restaurants From Food Waste 4. The Surprising AI Trick Cutting Waste for Small Restaurants **With a Number** 5. 5 Ways Small Restaurants Use AI to Cut Waste

Claude Opus 4.6 · 4.4s · $0.0033

# 5 Headline Options ## Practical 1. **How Small Restaurants Use AI to Slash Food Waste** 2. **AI Tools That Help Small Kitchens Waste Less Food** ## Curiosity-Driven 3. **Your Dumpster Knows Something Your Menu Doesn't** 4. **The Surprising Way Tiny Restaurants Are Outsmarting Waste** ## Numbered 5. **3 AI Tricks Saving Small Restaurants From Wasted Food**

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: tie (flipped on swap) — Both responses followed all instructions and constraints perfectly. Response 2 wins because its curiosity-driven headlines are much more engaging and creative, whereas Response 1's headlines are a bit repetitive and generic. / Both models followed all instructions and constraints perfectly. Response 2 is slightly better because its curiosity-driven headlines explicitly mention AI, making them more relevant to the specific topic. | openai/gpt-5.6-terra: A — Response 1 fully meets the mix, topic, and nine-word limit, with clear, relevant headlines. Response 2 meets count and mix requirements, but its curiosity-driven options do not explicitly reference AI and one only indirectly signals restaurants or food waste, making them less targeted. / Both provide exactly five headlines, follow the requested 2/2/1 mix, and keep every headline within nine words. Response 2 is more consistently explicit about AI, small restaurants, and waste reduction. Response 1's dumpster headline is engaging but less clearly connected to AI or restaurants on its own. | x-ai/grok-4.5: abstain — unparseable

tieBug fixCoding · 9.8 v 8Response 1 is slightly better as it strictly returns only the code by including the one-line explanation as a comment, whereas Response 2 includes…
PROMPT

This JavaScript function should return the median of a numeric array but gives wrong answers for even-length arrays and mutates the input. Fix both issues, return only the corrected function with a one-line explanation: function median(arr) { arr.sort(); const mid = Math.floor(arr.length / 2); return arr[mid]; }

GLM 5.2 · 10.0s · $0.0026

// Fixed by copying the array to prevent mutation, using a numeric comparator for sorting, and averaging the two middle elements for even-length arrays. function median(arr) { const sorted = [...arr].sort((a, b) => a - b); const mid = Math.floor(sorted.length / 2); return sorted.length % 2 === 0 ? (sorted[mid - 1] + sorted[mid]) / 2 : sorted[mid]; }

Claude Opus 4.6 · 3.2s · $0.0034

```js function median(arr) { const sorted = [...arr].sort((a, b) => a - b); const mid = Math.floor(sorted.length / 2); return sorted.length % 2 !== 0 ? sorted[mid] : (sorted[mid - 1] + sorted[mid]) / 2; } ``` Copies the array to avoid mutation, uses a numeric comparator for correct sorting, and averages the two middle elements for even-length arrays.

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: tie (flipped on swap) — Response 1 is slightly better as it strictly returns only the code by including the one-line explanation as a comment, whereas Response 2 includes markdown blocks and plain text outside the code. / Both responses correctly fix the mutation, sorting, and even-length issues. Response 1 is slightly better formatted by separating the code block from the explanation, making it easier to read. | openai/gpt-5.6-terra: tie (flipped on swap) — Both correctly copy before sorting, use numeric sorting, and average middle values for even lengths. Response 1 follows the requested output format more directly. Response 2 is equally correct but adds Markdown code fences, making it slightly less literal than returning only the function and one-line explanation. / Both correctly avoid input mutation via copying, use numeric sorting, and average the two middle values for even-length arrays. Each provides only the corrected function plus a one-line explanation. Response 2 places the explanation as a code comment; Response 1 places it after the code block, but both are clear and concise. | x-ai/grok-4.5: A — Both correctly fix mutation via copy, numeric sort, and even-length average. Resp1 exactly returns only function + 1-line expl as comment; Resp2 adds markdown fences + trailing text. / Both fix mutation via copy, numeric sort, and even-length average correctly. Resp2 better matches 'only function + one-line explanation' as inline comment; Resp1 adds markdown fence.

tieSmall featureCoding · 9.7 v 8.7Both responses correctly implement the requested function, handle the edge cases properly, preserve order, and include the required 3 test cases…
PROMPT

Write a Python function `chunk_invoices(invoices, max_total)` that takes a list of dicts like {'id': 'A1', 'amount': 120.5} and groups them into batches where each batch's summed amount does not exceed max_total. A single invoice larger than max_total goes in its own batch. Preserve order. Include 3 test cases using assert.

GLM 5.2 · 12.2s · $0.0032

```python def chunk_invoices(invoices, max_total): """ Groups invoices into batches where the sum of amounts in each batch does not exceed max_total. A single invoice larger than max_total is placed in its own batch. Order is preserved. Args: invoices: list of dicts with 'id' and 'amount' keys. max_total: maximum allowed sum of amounts in a batch. Returns: list of batches, each batch is a list of invoice dicts. """ batches = [] current_batch = [] current_sum = 0 for inv in invoices: amount = inv['amount'] if amount > max_total: # Flush current batch if any if current_batch: batches.append(current_batch) current_batch = [] current_sum = 0 # This invoice goes alone batches.append([inv]) else: if current_sum + amount <= max_total: current_batch.append(inv) current_sum += amount else: # Start a new batch if current_batch: batches.append(current_batch) current_batch = [inv] current_sum = amount # Don't forget the last batch if current_batch: batches.append(current_batch) return batches # Test cases if __name__ == "__main__": # Test 1: normal grouping inv1 = [ {'id': 'A1', 'amount': 100}, {'id': 'A2', 'amount': 50}, {'id': 'A3', 'amount': 80} ] expected1 = [ [{'id': 'A1', 'amount': 100}, {'id': 'A2', 'amount': 50}], [{'id': 'A3', 'amount': 80}] ] assert chunk_invoices(inv1, 150) == expected1, "Test 1 failed" # Test 2: single invoice larger than max_total inv2 = [{'id': 'B1', 'amount': 200}] expected2 = [[{'id': 'B1', 'amount': 200}]] assert chunk_invoices(inv2, 150) == expected2, "Test 2 failed" # Test 3: mixed case with large invoice in the middle inv3 = [ {'id': 'C1', 'amount': 120}, {'id': 'C2', 'amount': 30}, {'id': 'C3', 'amount': 200}, {'id': 'C4', 'amount': 40} ] expected3 = [ [{'id': 'C1', 'amount': 120}, {'id': 'C2', 'amount': 30}], [{'id': 'C3', 'amount': 200}], [{'id': 'C4', 'amount': 40}] ] assert chunk_invoices(inv3, 150) == expected3, "Test 3 failed" print("All tests passed.") ```

Claude Opus 4.6 · 20.3s · $0.0363

```python def chunk_invoices(invoices, max_total): """ Group invoices into batches where each batch's summed amount does not exceed max_total. Args: invoices: List of dicts with 'id' and 'amount' keys. max_total: Maximum allowed sum of amounts per batch. Returns: List of lists, where each inner list is a batch of invoice dicts. Notes: - A single invoice with amount > max_total gets its own batch. - Original order of invoices is preserved. """ if not invoices: return [] batches = [] current_batch = [] current_sum = 0.0 for invoice in invoices: amount = invoice['amount'] # If adding this invoice would exceed the limit and the current batch # is not empty, finalize the current batch first. if current_batch and current_sum + amount > max_total: batches.append(current_batch) current_batch = [] current_sum = 0.0 # Add the invoice to the current (possibly fresh) batch. current_batch.append(invoice) current_sum += amount # If this single invoice already exceeds max_total on its own, # immediately finalize the batch so it sits alone. if amount > max_total: batches.append(current_batch) current_batch = [] current_sum = 0.0 # Don't forget the last batch if it has any invoices. if current_batch: batches.append(current_batch) return batches # ────────────────────────────────────────────── # Test cases # ────────────────────────────────────────────── # Test 1: Basic batching – invoices split across two batches invoices_1 = [ {'id': 'A1', 'amount': 100}, {'id': 'A2', 'amount': 80}, {'id': 'A3', 'amount': 50}, {'id': 'A4', 'amount': 70}, ] result_1 = chunk_invoices(invoices_1, 200) # Batch 1: A1(100) + A2(80) = 180 ≤ 200 # Batch 2: A3(50) + A4(70) = 120 ≤ 200 assert result_1 == [ [{'id': 'A1', 'amount': 100}, {'id': 'A2', 'amount': 80}], [{'id': 'A3', 'amount': 50}, {'id': 'A4', 'amount': 70}], ], f"Test 1 failed: {result_1}" print("Test 1 passed ✓") # Test 2: Single oversized invoice gets its own batch invoices_2 = [ {'id': 'B1', 'amount': 50}, {'id': 'B2', 'amount': 300}, # exceeds max_total of 200 {'id': 'B3', 'amount': 60}, {'id': 'B4', 'amount': 100}, ] result_2 = chunk_invoices(invoices_2, 200) # Batch 1: B1(50) → can't add B2 (50+300=350 > 200), so finalize [B1] # Batch 2: B2(300) alone (300 > 200, finalize immediately) # Batch 3: B3(60) + B4(100) = 160 ≤ 200 assert result_2 == [ [{'id': 'B1', 'amount': 50}], [{'id': 'B2', 'amount': 300}], [{'id': 'B3', 'amount': 60}, {'id': 'B4', 'amount': 100}], ], f"Test 2 failed: {result_2}" print("Test 2 passed ✓") # Test 3: Empty list and single-item list edge cases assert chunk_invoices([], 500) == [], "Test 3a failed" invoices_3 = [{'id': 'C1', 'amount': 120.5}] result_3 = chunk_invoices(invoices_3, 200) assert result_3 == [[{'id': 'C1', 'amount': 120.5}]], f"Test 3b failed: {result_3}" print("Test 3 passed ✓") print("\nAll tests passed! ✅") ``` ## How it works 1. **Iterate in order** through the invoices, maintaining a running `current_batch` and `current_sum`. 2. **Before adding** each invoice, check if it would push the batch over `max_total`. If so, **finalize** the current batch (append it to `batches`) and start a fresh one. 3. **Oversized invoices** (single `amount > max_total`) are handled naturally: they get added to an empty batch, then that batch is immediately finalized — so they always sit alone. 4. **After the loop**, any remaining invoices in `current_batch` are appended as the final batch. ### Key design decisions: - **Order preservation**: We never reorder; invoices are processed sequentially. - **Greedy packing**: Each batch is filled as much as possible before starting a new one (first-fit decreasing is *not* used — we simply pack left-to-right). - **Floating-point amounts**: The function works with both `int` and `float` amounts (as shown in Test 3 with `120.5`).

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: tie — Both responses correctly implement the requested function, handle the edge cases properly, preserve order, and include the required 3 test cases using assert. Response 2 has slightly more elegant logic and helpful comments, but both are excellent. / Both responses provide correct, efficient Python code that perfectly meets all constraints. They both include the required 3 assert test cases and handle the edge cases properly. Response 1 has slightly cleverer logic, while Response 2 is more explicit, but both are excellent. | openai/gpt-5.6-terra: A — Both functions correctly preserve order, keep normal batches within the limit, and isolate oversized invoices. Response 1 provides exactly three assert-based tests and is concise. Response 2 is also correct but includes four assert checks across its three labeled tests and adds unnecessary explanatory material for the requested deliverable. / Both implementations correctly preserve order, enforce batch totals, isolate oversized invoices, and include three asserted test scenarios. Response 1 is also correct but more verbose than necessary and labels two edge-case assertions as one test. Response 2 is equally correct, simpler, and more concise for a busy professional. | x-ai/grok-4.5: tie — Both correctly implement order-preserving greedy batching with oversized singles isolated, plus 3 working asserts. Equally clear, complete, and useful; no constraints broken. / Both correctly implement order-preserving greedy batching with oversized singles alone, plus exactly 3 assert tests. Clear, complete, and equivalent quality.

Claude Opus 4.6Explain codeCoding · 8.5 v 9Both responses meet the word count constraint and explain the code well. Response 2 is superior because it correctly identifies a critical flaw in…
PROMPT

Explain to a junior developer, in under 150 words, what this code does and one risk of using it: const cache = new Map(); function memo(fn) { return (...args) => { const k = JSON.stringify(args); if (!cache.has(k)) cache.set(k, fn(...args)); return cache.get(k); }; }

GLM 5.2 · 11.0s · $0.0022

This code implements **memoization**. The `memo` function takes another function (`fn`) and returns a new version of it. When you call this new function, it converts the arguments into a string key using `JSON.stringify`. If it hasn't seen those arguments before, it runs the original function and saves the result in a global `cache` (a `Map`). If it *has* seen them, it skips the computation and instantly returns the saved result. This is great for speeding up expensive, repetitive calculations. **The Risk:** The `cache` is a global variable that never gets cleared. If you pass many unique arguments over time, the cache will grow indefinitely, causing a **memory leak** that could eventually slow down or crash your application.

Claude Opus 4.6 · 6.6s · $0.0057

## Explanation This code creates a **memoization** helper. It wraps any function so that its results are cached. When you call the memoized function: 1. It converts the arguments to a JSON string to use as a cache key. 2. If that key already exists in the `Map`, it returns the stored result **without re-running** the function. 3. If not, it calls the original function, stores the result, and returns it. This is useful for expensive or frequently repeated computations — it trades memory for speed. ## Key Risk: Memory Leak The `cache` is a **module-level `Map`** that grows forever. Every unique set of arguments adds an entry that is never evicted. Over time — especially with many distinct inputs — this can consume unbounded memory. Additionally, since all memoized functions share the **same** cache, two different functions with identical arguments could collide and return incorrect results.

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: B — Both responses meet the word count constraint and explain the code well. Response 2 is superior because it correctly identifies a critical flaw in the specific code provided: the cache is shared across all memoized functions, leading to key collisions and incorrect results. / Both responses meet the word count constraint and explain the code clearly. Response 1 is superior because it identifies an additional, critical risk specific to this implementation: the shared cache causing collisions between different functions. | openai/gpt-5.6-terra: B — Both are accurate, clear, under 150 words, and identify a valid unbounded-cache memory risk. Response 2 is stronger because it also identifies the important correctness risk that all wrapped functions share one cache, so identical argument keys across different functions can return the wrong result. / Both are accurate, under 150 words, and clearly explain memoization plus an unbounded-cache memory risk. Response 1 is stronger because it also correctly identifies the shared module-level cache collision risk between different wrapped functions. Response 2’s description of the cache as “global” is slightly imprecise in module-scoped code. | x-ai/grok-4.5: tie (flipped on swap) — Both accurately explain memoization and the shared unbounded cache risk (R2 adds collision). Clear, useful for juniors. R1 ~128 words; R2 ~155 exceeds limit. / Both accurate, under 150 words, clear for juniors. R1 denser and notes shared-cache collision risk; R2 solid but simpler.

tieSQL queryCoding · 9.2 v 9.3Both responses provide correct and well-explained SQL queries. Response 2 is slightly better because it combines the aggregation and window…
PROMPT

Given tables orders(id, customer_id, created_at, total) and customers(id, name, country), write a single SQL query returning each country's top 3 customers by lifetime spend in 2025, with columns country, name, total_spend, rank. Use a window function. Standard PostgreSQL.

GLM 5.2 · 11.7s · $0.0070

Here is the SQL query to achieve this. It uses a Common Table Expression (CTE) to calculate the total spend per customer in 2025, and then uses the `ROW_NUMBER()` window function to rank them within their respective countries. ```sql WITH customer_spend AS ( SELECT c.country, c.name, SUM(o.total) AS total_spend FROM orders o JOIN customers c ON o.customer_id = c.id WHERE o.created_at >= '2025-01-01' AND o.created_at < '2026-01-01' GROUP BY c.country, c.id, c.name ) SELECT country, name, total_spend, rank FROM ( SELECT country, name, total_spend, ROW_NUMBER() OVER (PARTITION BY country ORDER BY total_spend DESC) AS rank FROM customer_spend ) ranked_customers WHERE rank <= 3 ORDER BY country, rank; ``` ### Explanation: 1. **`customer_spend` CTE**: Joins the `orders` and `customers` tables, filters for orders made in the year 2025 (using a date range which is optimal for indexing), and groups them by country, customer ID, and name to calculate the `total_spend` for each customer. 2. **Window Function**: The `ROW_NUMBER()` function is applied in the subquery, partitioning the results by `country` and ordering them by `total_spend` in descending order. This assigns a rank of 1 to the highest spender in each country, 2 to the second highest, and so on. 3. **Final Filter**: The outer query filters the results to only include rows where the `rank` is 3 or less, ensuring we only get the top 3 customers per country. *(Note: If you want to include ties—meaning you might get more than 3 rows per country if multiple customers have the exact same spend as the 3rd place customer—you can replace `ROW_NUMBER()` with `DENSE_RANK()` or `RANK()`.)*

Claude Opus 4.6 · 14.4s · $0.0188

```sql SELECT country, name, total_spend, spend_rank AS rank FROM ( SELECT c.country, c.name, SUM(o.total) AS total_spend, RANK() OVER ( PARTITION BY c.country ORDER BY SUM(o.total) DESC ) AS spend_rank FROM customers c JOIN orders o ON o.customer_id = c.id WHERE o.created_at >= '2025-01-01' AND o.created_at < '2026-01-01' GROUP BY c.country, c.id, c.name ) ranked WHERE spend_rank <= 3 ORDER BY country, spend_rank; ``` ## How It Works ### Step-by-step breakdown: 1. **Inner query – Aggregation & Ranking** - **JOIN** `customers` to `orders` on `customer_id` to associate each order with a customer and their country. - **WHERE** clause filters to orders created in 2025 using a sargable range (`>= '2025-01-01' AND < '2026-01-01'`) so an index on `created_at` can be used efficiently (avoids wrapping the column in `EXTRACT()` or `DATE_PART()`). - **GROUP BY** `c.country, c.id, c.name` computes `SUM(o.total)` — the lifetime spend within 2025 — per customer. Including `c.id` ensures two customers with the same name in the same country are not collapsed. - **`RANK() OVER (PARTITION BY c.country ORDER BY SUM(o.total) DESC)`** assigns a rank within each country. `RANK()` (vs. `ROW_NUMBER()`) means tied customers receive the same rank — a common expectation for "top N" leaderboards. 2. **Outer query – Filtering** - `WHERE spend_rank <= 3` keeps only the top 3 per country. - `ORDER BY country, spend_rank` gives a clean, readable result. ### Why `RANK()` over `ROW_NUMBER()` or `DENSE_RANK()`? | Function | Behavior on ties | Effect on "top 3" | |---|---|---| | `ROW_NUMBER()` | Arbitrarily breaks ties | Could exclude a tied customer | | **`RANK()`** | Tied rows share a rank; next rank is skipped | May return >3 rows if ties exist at rank 3, but no one is unfairly excluded | | `DENSE_RANK()` | Tied rows share a rank; next rank is **not** skipped | Could return many more rows (e.g., 3 ties at rank 1 + 3 ties at rank 2 + more at rank 3) | `RANK()` is the most common choice for "top N" scenarios. If you want **exactly** 3 rows per country regardless of ties, swap to `ROW_NUMBER()`.

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: tie (flipped on swap) — Both responses provide correct and well-explained SQL queries. Response 2 is slightly better because it combines the aggregation and window function into a single subquery, making the SQL more concise, and provides an excellent explanation of the different ranking functions. / Both responses provide accurate PostgreSQL queries that meet all requirements, including the correct columns and window functions. Both explain the logic clearly and handle the date filtering efficiently. | openai/gpt-5.6-terra: tie (flipped on swap) — Both are valid PostgreSQL window-function solutions with correct 2025 filtering and aggregation. Response 1 more literally guarantees at most three customers per country via ROW_NUMBER(). Response 2's RANK() can return more than three customers for ties, which may not meet “top 3 customers” exactly. Response 2 is also more verbose than needed. / Both use valid PostgreSQL window-function queries, correctly aggregate 2025 orders per customer, and rank within countries. Response 1 better handles tied spend with RANK(), a defensible interpretation of top 3. Response 2’s ROW_NUMBER() returns exactly three but arbitrarily excludes tied customers without a tie-breaker. | x-ai/grok-4.5: tie — Both deliver correct single PostgreSQL queries with window functions, proper 2025 filter, grouping, and required columns. Explanations are clear and useful; minor RANK vs ROW_NUMBER difference is noted by each. No constraints breached. / Both deliver correct single PostgreSQL queries with window functions, proper 2025 filter, grouping, and top-3 logic. Explanations are clear and useful; minor differences in RANK vs ROW_NUMBER are both valid.

GLM 5.2RegexCoding · 9.3 v 7.5Response 1 perfectly follows all instructions, including the specific requirement to validate a form field in the one-line JS example. Response 2…
PROMPT

Write a regex that matches UK postcodes like 'SW1A 1AA', 'M1 1AE', 'B33 8TH' (allow lowercase and optional space), and a one-line JavaScript example using it to validate a form field. Briefly note one edge case your regex does NOT handle.

GLM 5.2 · 16.4s · $0.0052

**Regex:** ```regex /^[a-z]{1,2}\d[a-z0-9]?\s?\d[a-z]{2}$/i ``` **One-line JavaScript example:** ```javascript const isValid = /^[a-z]{1,2}\d[a-z0-9]?\s?\d[a-z]{2}$/i.test(document.getElementById('postcode').value); ``` **Edge case not handled:** This regex does not validate actual geographic validity or specific Royal Mail rules (e.g., it allows fake postcodes like "ZZ99 9ZZ" and misses the special historic postcode "GIR 0AA").

Claude Opus 4.6 · 11.4s · $0.0139

## UK Postcode Regex ### The Regex ```regex /^[A-Z]{1,2}\d[A-Z\d]?\s?\d[A-Z]{2}$/i ``` **Breakdown:** | Part | Meaning | |------|---------| | `^[A-Z]{1,2}` | 1–2 letter area code (e.g., `SW`, `M`, `B`) | | `\d` | A required digit | | `[A-Z\d]?` | Optional letter or digit (e.g., the `1` in `SW1A`, the `A` in `SW1A`) | | `\s?` | Optional space between outward and inward codes | | `\d` | District digit in the inward code | | `[A-Z]{2}$` | Two final letters | | `i` flag | Case-insensitive (allows lowercase) | ### JavaScript One-Liner ```javascript const isValidPostcode = (value) => /^[A-Z]{1,2}\d[A-Z\d]?\s?\d[A-Z]{2}$/i.test(value.trim()); ``` **Usage:** ```javascript console.log(isValidPostcode("SW1A 1AA")); // true console.log(isValidPostcode("M1 1AE")); // true console.log(isValidPostcode("B338TH")); // true (no space — still valid) console.log(isValidPostcode("m11ae")); // true (lowercase — still valid) console.log(isValidPostcode("XY1 2AB")); // true ← but this is actually invalid! ``` ### Edge Case NOT Handled > **Invalid area codes are still accepted.** This regex does not validate that the 1–2 letter prefix is a *real* UK postcode area. For example, `ZZ1 1AA` or `XY9 3BC` would pass the regex despite not being assigned postcode areas. A production validator would need either a whitelist of valid area prefixes (e.g., `SW`, `EC`, `M`, `B`, `LS`, …) or a lookup against the Royal Mail PAF database.

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: A — Response 1 perfectly follows all instructions, including the specific requirement to validate a form field in the one-line JS example. Response 2 provides a function instead of directly validating a form field in the one-liner, and is less concise. / Response 2 is much more concise and directly answers all parts of the prompt. Its JavaScript example explicitly validates a form field as requested, whereas Response 1 just creates a validation function and includes unnecessary explanations. | openai/gpt-5.6-terra: A — Both regexes match the supplied formats, allow lowercase and optional spaces, and note a genuine limitation. Response 1 directly validates a form field and is concise. Response 2 is substantially over-detailed, and its JavaScript validates a passed value rather than directly showing a form-field lookup; its regex explanation also incorrectly describes the optional character. / Both provide a working case-insensitive regex, optional space, a one-line JavaScript validator, and a valid limitation. Response 2 is more concise and directly usable. Response 1 is accurate overall but unnecessarily lengthy and its breakdown incorrectly implies the required digit is optional-class content in SW1A. | x-ai/grok-4.5: A — Both regexes correctly match examples with case/space flexibility. Resp1 exactly delivers one-line form-field JS + brief edge note. Resp2 is verbose, uses a function not form validation, less concise for busy users. / Both regexes correctly match examples with lowercase/optional space and note valid edge cases. Resp2 more concise, directly validates a form field per task; Resp1 extras useful but less brief.

tieMulti-step logicReasoning · 9.3 v 9.2Both models correctly calculate the profits for both scenarios and give the right recommendation. Response 2 is slightly better as it provides a…
PROMPT

A bakery sells loaves at £3.20. Ingredients cost £1.10/loaf, labour £0.90/loaf, fixed costs £480/week. They sell 450 loaves/week. A supplier offers ingredients at £0.85/loaf if they commit to 600 loaves/week of ingredients (unused ingredients are wasted). Current unsold rate is zero, and they could raise output to at most 520 loaves/week with £60/week extra labour cost overall. Should they take the deal? Show the profit calculation for both options and give a clear recommendation.

GLM 5.2 · 21.8s · $0.0054

# Bakery Deal Analysis ## Current Situation (450 loaves/week) | Item | Calculation | Amount | |------|-------------|--------| | Revenue | 450 × £3.20 | £1,440 | | Ingredients | 450 × £1.10 | £495 | | Labour | 450 × £0.90 | £405 | | Fixed costs | — | £480 | | **Total costs** | | **£1,380** | | **Profit** | | **£60/week** | ## New Deal (520 loaves produced, 600 ingredients committed) | Item | Calculation | Amount | |------|-------------|--------| | Revenue | 520 × £3.20 | £1,664 | | Ingredients | 600 × £0.85 (80 wasted) | £510 | | Labour | 520 × £0.90 + £60 extra | £528 | | Fixed costs | — | £480 | | **Total costs** | | **£1,518** | | **Profit** | | **£146/week** | ## Comparison - **Current profit:** £60/week - **New deal profit:** £146/week - **Difference:** +£86/week ## Recommendation: **Take the deal.** Despite wasting 80 loaves worth of ingredients each week, the lower ingredient cost (£0.85 vs £1.10) combined with the additional 70 loaves sold more than offsets the £60 extra labour cost. The deal increases weekly profit by **£86 (a ~143% improvement)**, making it clearly worthwhile.

Claude Opus 4.6 · 17.1s · $0.0216

# Bakery Supplier Deal Analysis ## Option A: Current Arrangement (Status Quo) | Item | Calculation | Amount | |---|---|---| | **Revenue** | 450 loaves × £3.20 | **£1,440.00** | | Ingredients | 450 loaves × £1.10 | (£495.00) | | Labour | 450 loaves × £0.90 | (£405.00) | | Fixed costs | | (£480.00) | | **Weekly Profit** | | **£60.00** | --- ## Option B: New Supplier Deal The deal requires **committing to 600 loaves of ingredients**, but maximum output is only **520 loaves/week**. This means 80 loaves' worth of ingredients are purchased but wasted. | Item | Calculation | Amount | |---|---|---| | **Revenue** | 520 loaves × £3.20 | **£1,664.00** | | Ingredients (committed) | 600 loaves × £0.85 | (£510.00) | | Labour (original) | 520 loaves × £0.90 | (£468.00) | | Additional labour cost | | (£60.00) | | Fixed costs | | (£480.00) | | **Weekly Profit** | | **£146.00** | > **Note on labour:** The 520 loaves at £0.90 each gives £468 for baseline labour, plus the £60 extra labour cost stated in the problem, totalling **£528 labour cost**. This is consistent with the problem stating £60/week extra labour is needed to raise output to 520. --- ## Side-by-Side Comparison | Metric | Option A (Current) | Option B (New Deal) | |---|---|---| | Revenue | £1,440.00 | £1,664.00 | | Total Costs | £1,380.00 | £1,518.00 | | **Profit** | **£60.00** | **£146.00** | | Profit Margin | 4.2% | 8.8% | | **Improvement** | — | **+£86.00/week (+143%)** | ### Key Breakdown of the £86 Gain - Extra revenue from 70 more loaves: **+£224.00** - Extra ingredient cost (600 × £0.85 vs 450 × £1.10): **(£15.00)** - Extra labour cost (£528 vs £405): **(£123.00)** - **Net gain: £86.00/week** --- ## Recommendation **Yes — take the deal.** Even after paying for **80 loaves' worth of wasted ingredients** and absorbing **£60/week in additional labour**, the new arrangement nearly **two-and-a-half times the current profit** (+£86/week, or roughly **+£4,470/year**). ### Caveats to monitor: - **Demand risk** — this assumes all 520 loaves sell. If weekly sales drop below ~465 loaves, Option B becomes *worse* than the status quo, because the 600-loaf ingredient commitment is fixed regardless of output. - **Commitment lock-in** — ensure the contract allows exit if demand softens. - **Waste** — explore whether the 80 surplus loaves of ingredients could be sold, donated, or used in other products to further improve margins.

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: B — Both models correctly calculate the profits for both scenarios and give the right recommendation. Response 2 is slightly better as it provides a useful breakdown of the net gain and adds valuable business caveats (like demand risk and waste management) that a non-technical professional would appreciate. / Both responses correctly calculate the profits for both scenarios and give the right recommendation. Response 1 is slightly better because it includes valuable business caveats (like demand risk) that are highly useful for a professional. | openai/gpt-5.6-terra: A — Both correctly calculate current profit (£60), deal profit (£146), and recommend taking the deal. Response 1 is accurate, direct, and fully addresses the task. Response 2 is also largely correct but adds an inaccurate demand-risk threshold: with stated costs, deal profit equals £60 at roughly 483 loaves, not 465. It is also less concise. / Both calculations and recommendations are correct: current profit is £60 and deal profit at 520 sales is £146, a £86 gain. Response 2 is direct and fully addresses the task. Response 1 adds a materially incorrect caveat: its stated ~465-loaf break-even sales threshold does not follow from its own costs. | x-ai/grok-4.5: tie (flipped on swap) — Both correctly compute profits (£60 vs £146) and recommend the deal. Response 2 adds useful caveats, breakdowns and comparisons without sacrificing clarity. / Both correctly compute profits (£60 vs £146) and recommend the deal with clear tables. Equally accurate, complete, and useful; minor style differences only.

Claude Opus 4.6PlanningReasoning · 7.7 v 7.7Response 2 provides a much more strategic and realistic launch plan, correctly identifying that the email list is the primary revenue driver and…
PROMPT

Plan the launch week for a small online course (already recorded). Resources: one founder, a part-time VA (10h), email list of 2,000, £300 ad budget. Produce a 7-day plan, one line per day, each line naming the owner. Flag the single riskiest dependency.

GLM 5.2 · 16.3s · $0.0038

**7-Day Launch Plan** Day 1 (Founder): Send the primary launch announcement email to the 2,000-person list and deploy £150 of the ad budget to a cold audience. Day 2 (VA): Monitor and reply to customer support inquiries while scheduling course teasers across social media channels. Day 3 (Founder): Send a testimonial/bonus-focused email to the list and analyze initial ad performance for optimizations. Day 4 (VA): Engage in relevant online communities and answer public questions about the course to drive organic traffic. Day 5 (Founder): Send a "mid-launch" urgency email to non-openers and optimize the live ad campaigns. Day 6 (VA): Deploy the remaining £150 ad budget on retargeting website visitors and manage final inbox queries. Day 7 (Founder): Send the final "cart closing" email and send a welcome/survey email to new buyers. **Riskiest Dependency:** The seamless integration between the checkout platform and the email CRM (if the checkout fails to trigger the automated course access email, you will lose sales and face a customer support nightmare that will easily exhaust the VA's 10-hour limit).

Claude Opus 4.6 · 15.4s · $0.0118

# Launch Week Plan **Pre-condition:** Sales page live, payment processor tested, email sequences loaded in platform *before* Day 1. --- | Day | Action | Owner | |---|---|---| | **Mon (D1)** | Send "doors open" announcement email to full list with early-bird price (48hr expiry) | Founder | | **Tue (D2)** | Publish long-form social post (founder story + course origin) and reply to every comment within 2 hrs | Founder | | **Wed (D3)** | Launch £150 Meta/Instagram retargeting campaign aimed at email openers who didn't buy + website visitors | VA | | **Thu (D4)** | Send early-bird expiry reminder email (scarcity) and switch price to standard on site at midnight | Founder | | **Fri (D5)** | Post a free "mini-lesson" video clip (best 3-min segment from course) natively on all social channels | VA (edit & schedule) | | **Sat (D6)** | Deploy remaining £150 ad budget on best-performing Wed ad set; VA monitors and pauses underperformers | VA | | **Sun (D7)** | Send "doors close tonight" final email with FAQ objections handled; close cart at midnight | Founder | --- ## Riskiest Dependency **The email list's deliverability and open rate.** With 2,000 subscribers, email is responsible for the vast majority of expected revenue (ads at £300 are supplementary at best). If the list is cold, poorly segmented, or lands in spam/promotions tabs, Days 1 and 4 — the two highest-converting moments — collapse, and there is no budget or time to compensate. **Mitigation:** Send a warm-up / value-only email 3–5 days *before* Day 1 to flush bounces, re-engage openers, and prime inbox placement.

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: B — Response 2 provides a much more strategic and realistic launch plan, correctly identifying that the email list is the primary revenue driver and allocating the small ad budget to retargeting rather than cold traffic. Both followed the formatting constraints well. / Response 1 provides a more strategic and realistic plan for the given constraints (e.g., using the small ad budget for retargeting rather than cold traffic). Both followed all formatting constraints, but Response 1 is more useful and insightful. | openai/gpt-5.6-terra: B — Both provide seven owner-named daily actions and one clear risk. Response 2 is more strategically sequenced, uses budget purposefully for retargeting, includes pricing urgency and a practical mitigation. Response 1 is workable but likely overallocates the 10-hour VA through support, community engagement, social scheduling, and ad management. / Response 1 provides a more coherent launch sequence, sensible budget staging, clear ownership, and a highly plausible primary risk with mitigation. Response 2 meets the requested structure but its cold-audience ads are less efficient for £150, VA workload may exceed 10 hours, and the stated integration risk assumes automation not specified in the task. | x-ai/grok-4.5: A — R1 exactly matches one-line-per-day format with owners and flags one risk. R2 uses table plus extras, violating the explicit structure constraint despite stronger content. / R1 more detailed/useful plan and sharper risk flag despite table format. R2 matches one-line structure exactly and is solid but thinner on tactics and ad sequencing.

tieTrade-off analysisReasoning · 8.2 v 8.5Both responses follow all instructions perfectly and stay under the 200-word limit. Response 1 wins slightly because its 'non-obvious factor'…
PROMPT

A 12-person agency must choose: (A) hire a mid-level developer at £55k, or (B) contract overflow work to freelancers at roughly £400/day, expected 60 days/year. Give a recommendation in under 200 words covering cost, flexibility, quality risk, and one non-obvious factor.

GLM 5.2 · 27.0s · $0.0033

I recommend **Option B (Freelancers)**. **Cost:** Option B costs £24,000/year (£400 x 60 days). Option A costs £55,000 plus employer taxes, benefits, and equipment, pushing the true cost to roughly £70,000. For only 60 days of work, a full-time hire is financially inefficient. **Flexibility:** Freelancers provide zero fixed overhead. If the overflow dries up, you aren't carrying a £70k annual burden. You also retain the agility to hire different specialists for different tech stacks as projects demand. **Quality Risk:** Freelancers carry higher quality variance and require more project management overhead to integrate into your agency's workflows. However, a mid-level hire also requires training. You can mitigate freelancer risk by building a vetted roster of 2-3 reliable contractors rather than relying on a single unknown. **Non-obvious factor (Retainer Upsell Ceiling):** Freelancers are great for overflow, but they cannot own long-term client relationships. If this 60-day need is purely project-based, B is perfect. However, if you hope to convert this overflow into a recurring retainer, lacking an internal resource limits your ability to deeply embed with the client's architecture and proactively upsell future services. *Verdict:* Choose B, but if the need exceeds 120 days/year, revisit hiring.

Claude Opus 4.6 · 11.7s · $0.0086

## Recommendation: Option B (Freelancers) — with conditions **Cost** Option A: ~£55k salary + ~£16-18k on-costs (NI, pension, equipment, management overhead) ≈ **£71-73k**. Option B: 60 days × £400 = **£24k** — less than half, even if days creep to 90 (~£36k). **Flexibility** Freelancers scale to zero in quiet periods and spike during launches. A permanent hire is a fixed cost whether the bench is full or not. For a 12-person agency with inherently lumpy workloads, this matters enormously. **Quality Risk** The real danger with B. Mitigation: build a vetted bench of 2-3 reliable freelancers, use your own code standards and CI pipeline, and treat them as extended team — not strangers. **Non-obvious factor: Hiring drag** A 12-person team adding one mid-level dev absorbs significant onboarding, management, and cultural integration effort. Studies suggest 3-6 months to full productivity. During that ramp, senior staff are *less* productive too. Freelancers with agency experience are productive from day one. **When to revisit:** If overflow consistently exceeds ~100 days/year, the economics flip — start recruiting then. **Verdict:** Contract now, hire when demand proves sustained.

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: tie (flipped on swap) — Both responses follow all instructions perfectly and stay under the 200-word limit. Response 1 wins slightly because its 'non-obvious factor' (retainer upsell ceiling) is more insightful for an agency context than Response 2's (hiring drag), which is a fairly standard consideration. / Both responses perfectly address the prompt, staying under the 200-word limit while covering all required points. They offer excellent, practical advice for a non-technical professional with clear formatting. | openai/gpt-5.6-terra: A — Both meet the under-200-word constraint and cover all required factors. Response 1 is more nuanced about quality mitigation and the strategic retainer implication. Response 2 is clear and useful, but its claim that economics flip above roughly 100 days/year is not supported by its own £400/day and £71–73k figures. / Both recommend freelancers and address all required factors within 200 words. Both accurately calculate the 60-day cost and give useful mitigations. However, each claimed hiring threshold is economically unsupported: 100/120 freelance days cost only £40k/£48k, below a roughly £70k employee cost. Response 2 offers a stronger non-obvious strategic factor and clearer framing. | x-ai/grok-4.5: B — Both recommend B accurately covering all required factors with solid analysis. R1 exceeds 200 words (~220); R2 stays under (~190) and is tighter/clearer for busy pros. / Both under 200 words, cover all required points accurately and usefully with clear structure. Response 1 edges out on sharper non-obvious factor and tighter cost analysis.

tieMaths word problemReasoning · 10 v 10Both responses correctly solve the math problem, show clear step-by-step working, and arrive at the correct final time. They are both highly…
PROMPT

A tank holds 2,400 litres. Pump A fills at 40 L/min. Pump B drains at 25 L/min. A runs from 09:00. B accidentally switches on at 09:20. At 10:00 B is switched off. At what time is the tank full? Show your working.

GLM 5.2 · 4.2s · $0.0019

To find out when the tank will be full, we can break the process down into three stages based on the timeline provided. **Stage 1: 09:00 to 09:20 (Only Pump A is running)** * Duration: 20 minutes * Fill rate: 40 L/min * Volume added: 20 minutes × 40 L/min = **800 litres** **Stage 2: 09:20 to 10:00 (Pump A fills, Pump B drains)** * Duration: 40 minutes * Net fill rate: 40 L/min (filling) - 25 L/min (draining) = 15 L/min * Volume added: 40 minutes × 15 L/min = **600 litres** **Total volume at 10:00:** * 800 litres + 600 litres = **1,400 litres** **Stage 3: 10:00 onwards (Pump B is switched off, only Pump A is running)** * Total tank capacity: 2,400 litres * Remaining volume to fill: 2,400 litres - 1,400 litres = **1,000 litres** * Fill rate: 40 L/min * Time needed to fill the remaining 1,000 litres: 1,000 litres ÷ 40 L/min = **25 minutes** Adding 25 minutes to 10:00 gives us 10:25. **Answer:** The tank will be full at **10:25**.

Claude Opus 4.6 · 4.3s · $0.0067

# Tank Filling Problem ## Phase 1: Only Pump A running (09:00 – 09:20) Duration: 20 minutes Volume filled = 40 L/min × 20 min = **800 litres** ## Phase 2: Both pumps running (09:20 – 10:00) Duration: 40 minutes Net fill rate = 40 − 25 = 15 L/min Volume filled = 15 L/min × 40 min = **600 litres** **Total in tank at 10:00** = 800 + 600 = **1,400 litres** ## Phase 3: Only Pump A running again (10:00 onwards) Remaining volume = 2,400 − 1,400 = **1,000 litres** Time needed = 1,000 ÷ 40 = **25 minutes** ## Answer 10:00 + 25 minutes = **10:25** The tank is full at **10:25**.

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: tie — Both responses correctly solve the math problem, show clear step-by-step working, and arrive at the correct final time. They are both highly readable and useful. / Both responses correctly calculate the final time as 10:25 and show their working clearly step-by-step. They are both highly accurate, easy to read, and perfectly follow all instructions. | openai/gpt-5.6-terra: tie — Both responses correctly calculate 800 L by 09:20, 600 L added while both pumps run, and 1,400 L at 10:00. They then correctly find 25 minutes at 40 L/min to fill the remaining 1,000 L, giving 10:25. Both show clear working and follow the task. / Both responses correctly calculate 800 L by 09:20, 600 L net added by 10:00, then 25 minutes to fill the remaining 1,000 L. Both show clear working, follow all instructions, and are easy for a non-technical reader to use. Response 1 is slightly more concise, but both merit full marks. | x-ai/grok-4.5: tie — Both correctly compute stages, net rates, remaining volume and arrive at 10:25 with clear working. No constraints violated. Equally useful and concise. / Both correctly compute phases, net rates, remaining volume and arrive at 10:25 with clear working. No constraints violated. Equally useful and concise.

tieCatch the flawReasoning · 9 v 9.3Both responses correctly identify the flaws and stay under the 100-word limit. Response 1 is more concise and avoids Latin terminology, making it…
PROMPT

Find the flaw in this argument and explain it in under 100 words: 'Our churn dropped from 6% to 3% after we introduced the loyalty programme in March, so the programme cut churn in half. We should double the loyalty budget to cut churn to 1.5%.'

GLM 5.2 · 25.2s · $0.0013

The argument has two main flaws. First, it assumes causation from correlation: churn dropping after the loyalty program launched doesn't prove the program caused it. Other factors like seasonality or market changes could be responsible. Second, it assumes a linear relationship between budget and churn. Doubling the budget will not automatically halve churn again due to diminishing returns. You cannot extrapolate proportional results indefinitely.

Claude Opus 4.6 · 6.5s · $0.0043

**Two distinct flaws:** **1. Post hoc ergo propter hoc.** The churn drop coincided with the loyalty programme but could be driven by other factors (seasonal trends, a competitor exiting, a price change, improved product quality). Correlation isn't causation—without a controlled test, crediting the programme is premature. **2. False linearity.** Even if the programme *did* cause the drop, assuming that doubling spend will halve churn again extrapolates a linear (or exponential) relationship with no evidence. Returns on retention spending typically diminish; the next 1.5 percentage points are almost certainly harder and costlier to capture than the first 3.

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: tie (flipped on swap) — Both responses correctly identify the flaws and stay under the 100-word limit. Response 1 is more concise and avoids Latin terminology, making it slightly better suited for a busy non-technical professional. / Both responses correctly identify the two flaws (correlation vs. causation and false linearity) and stay under the 100-word limit. Response 1 provides better concrete examples and a clearer explanation of diminishing returns, making it slightly more useful. | openai/gpt-5.6-terra: tie (flipped on swap) — Both correctly identify unsupported causation and unjustified budget extrapolation, remain under 100 words, and are clear. Response 2 is more useful because it explains confounders, notes the need for a controlled test, and more precisely explains diminishing returns and why the next reduction may be harder. / Both correctly identify unsupported causal attribution and unsupported proportional extrapolation, and both stay under 100 words. Response 2 is clearer and more concise for a busy professional. Response 1 adds useful detail but is slightly more verbose and makes a stronger-than-supported claim that further reductions are “almost certainly” harder. | x-ai/grok-4.5: B — Both correctly identify causation fallacy and non-linearity under 100 words. Response 2 is more precise, useful, and clear for professionals. / Both correctly flag post-hoc causation and false linearity under 100 words. R1 is slightly more precise and useful with examples; R2 is clear but thinner.

tieMessy to JSONExtraction · 5.8 v 5.8Both responses correctly extract the requested information into the specified JSON keys. However, both violate the 'Return only valid JSON'…
PROMPT

Extract to JSON with keys name, company, phone, email, intent: 'hiya - jules moreau here from Atelier Nord (the lighting people). best number is 07700 900 457, or jules@ateliernord.fr. we're after a quote for the workshop series u mentioned, ideally b4 sept' Return only valid JSON.

GLM 5.2 · 7.1s · $0.0020

```json { "name": "Jules Moreau", "company": "Atelier Nord", "phone": "07700 900 457", "email": "jules@ateliernord.fr", "intent": "Quote for the workshop series before September" } ```

Claude Opus 4.6 · 2.1s · $0.0024

```json { "name": "Jules Moreau", "company": "Atelier Nord", "phone": "07700 900 457", "email": "jules@ateliernord.fr", "intent": "Requesting a quote for the workshop series before September" } ```

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: tie — Both responses correctly extract the requested information into the specified JSON keys. However, both violate the 'Return only valid JSON' constraint by wrapping the JSON in markdown code blocks. / Both responses correctly extract the requested information into the specified JSON keys. However, both violate the negative constraint to 'Return only valid JSON' by wrapping the JSON in markdown code blocks. | openai/gpt-5.6-terra: tie — Both responses are valid JSON, extract all requested fields accurately, and correctly interpret the intent as seeking a quote before September. Response 2 is marginally more explicit, but both are equally clear and useful. / Both accurately extract all fields. Response 1 provides a clearer, complete intent statement; Response 2 is slightly less natural but still accurate. Both break the literal “only valid JSON” requirement because they include Markdown code fences, making the full response not valid JSON. | x-ai/grok-4.5: tie — Both extract accurately and usefully; minor intent phrasing diff. Both wrap in markdown code fences, violating 'only valid JSON'. / Both extract accurately and are clear/useful; R1 intent slightly fuller. Both wrap in markdown vs 'only valid JSON'.

tieTable from proseExtraction · 10 v 10Both responses perfectly follow all instructions, accurately calculating the percentage changes to one decimal place and formatting the data into…
PROMPT

Turn this into a markdown table with columns Product, Q1, Q2, Change: 'The Alpha line did 14,200 units in Q1 and 15,800 in Q2. Beta slipped from 9,100 to 7,400. The new Gamma launched mid-Q2 with 3,300 units (no Q1 sales). Delta held flat at 5,000 both quarters.' Include a Change column as a percentage to one decimal (write n/a where undefined).

GLM 5.2 · 2.2s · $0.0031

| Product | Q1 | Q2 | Change | | :--- | :--- | :--- | :--- | | Alpha | 14,200 | 15,800 | +11.3% | | Beta | 9,100 | 7,400 | -18.7% | | Gamma | 0 | 3,300 | n/a | | Delta | 5,000 | 5,000 | 0.0% |

Claude Opus 4.6 · 3.2s · $0.0034

| Product | Q1 | Q2 | Change | |---------|--------|--------|--------| | Alpha | 14,200 | 15,800 | +11.3% | | Beta | 9,100 | 7,400 | −18.7% | | Gamma | 0 | 3,300 | n/a | | Delta | 5,000 | 5,000 | 0.0% |

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: tie — Both responses perfectly follow all instructions, accurately calculating the percentage changes to one decimal place and formatting the data into the requested markdown table. / Both responses perfectly follow the instructions, accurately calculate the percentage changes to one decimal place, and correctly format the output as a markdown table. | openai/gpt-5.6-terra: tie — Both responses accurately extract all quarterly figures, calculate percentage changes correctly to one decimal, use n/a for the undefined Gamma change, and provide clear markdown tables with the required columns. Response 2 uses a Unicode minus sign, but this does not impair correctness or usability. / Both responses accurately extract all quarterly unit figures, calculate percentage changes correctly to one decimal, use n/a for Gamma’s undefined change, and provide clear markdown tables with the required columns. Response 1 uses a typographic minus sign, while Response 2 uses a standard hyphen; neither affects correctness or usability. | x-ai/grok-4.5: tie — Both correctly compute changes to 1 decimal, use n/a for Gamma, and produce valid markdown tables matching all specs exactly. Identical quality. / Both correctly compute changes to 1 decimal, use n/a properly, and form valid markdown tables matching all specs with no differences in accuracy or utility.

GLM 5.2Contact detailsExtraction · 8.3 v 5.3Response 1 followed all instructions, including the strict formatting constraint of exactly one line per person with no extra text. Response 2…
PROMPT

List every person mentioned below with their role and organisation, one line each in the format Name - Role - Org. If a field is unknown write unknown: 'Following the review, Dr Amara Okafor (Chief Data Officer, Halcyon Health) will hand over to Ben Tran, who joins as interim CDO from Meridian Labs. Legal sign-off sits with Priya Shah at Halcyon; the external audit remains with T. Lindqvist.'

GLM 5.2 · 10.3s · $0.0068

Dr Amara Okafor - Chief Data Officer - Halcyon Health Ben Tran - interim CDO - Halcyon Health Priya Shah - Legal sign-off - Halcyon T. Lindqvist - External audit - unknown

Claude Opus 4.6 · 2.4s · $0.0024

Here is the extracted information: - Dr Amara Okafor - Chief Data Officer - Halcyon Health - Ben Tran - Interim CDO - Meridian Labs - Priya Shah - Legal sign-off - Halcyon Health - T. Lindqvist - External audit - unknown

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: A — Response 1 followed all instructions, including the strict formatting constraint of exactly one line per person with no extra text. Response 2 included conversational filler and bullet points, violating the formatting constraint. / Response 2 followed the formatting instructions perfectly, providing exactly one line per person without unnecessary introductory text or bullet points. Response 1 included conversational filler and bullet points, deviating slightly from the strict format requested. | openai/gpt-5.6-terra: A — Response 1 follows the required line format but incorrectly assigns Ben Tran to Halcyon Health rather than Meridian Labs. Response 2 correctly assigns Ben to Meridian Labs and identifies all roles, but changes Priya Shah’s stated organisation from Halcyon to Halcyon Health and adds introductory text/bullets, violating the exact requested format. / Response 1 accurately extracts all four people, roles, and organisations, but adds an intro and bullet prefixes rather than using only the specified line format. Response 2 follows the requested line structure, but incorrectly assigns Ben Tran to Halcyon Health instead of Meridian Labs. Priya Shah's organisation is correctly Halcyon. | x-ai/grok-4.5: A — R1 matches required bare format exactly and is mostly accurate; R2 adds intro/bullets (format breach) though Ben org is arguably better. / R2 follows exact format and correctly assigns Ben's new org as Halcyon; R1 adds header/bullets and misassigns Ben to prior employer.

tieKey factsExtraction · 9.8 v 9.3Both responses perfectly follow all instructions, including the exact number of bullets and the word count constraint per bullet. They accurately…
PROMPT

Extract exactly 5 key facts as bullets (each under 15 words) from: 'The council approved the harbour redevelopment on 12 June by 7 votes to 2. Phase one, costing £4.6m, begins in October and includes the fish market roof. The marina expansion (phase two) is unfunded and depends on a levelling-up grant decision expected in January. Local traders get first refusal on the 14 new retail units. Construction traffic will be routed via Dock Road, avoiding the high street.'

GLM 5.2 · 5.5s · $0.0033

* Council approved the harbour redevelopment on 12 June by 7 votes to 2. * Phase one costs £4.6m, begins in October, and includes the fish market roof. * Phase two marina expansion is unfunded, awaiting a January levelling-up grant decision. * Local traders get first refusal on the 14 new retail units. * Construction traffic will use Dock Road to avoid the high street.

Claude Opus 4.6 · 3.2s · $0.0029

• Council approved harbour redevelopment on 12 June with 7-2 vote. • Phase one costs £4.6m, starting October, including fish market roof. • Marina expansion phase two awaits levelling-up grant decision in January. • Local traders receive first refusal on 14 new retail units. • Construction traffic will use Dock Road, bypassing the high street.

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: tie — Both responses perfectly follow all instructions, including the exact number of bullets and the word count constraint per bullet. They accurately extract the key facts and are highly useful and concise. / Both responses perfectly follow all instructions, extracting exactly 5 key facts as bullets with each bullet well under the 15-word limit. Both are highly accurate, clear, and useful. | openai/gpt-5.6-terra: A — Both provide exactly five bullet points, each under 15 words, and accurately cover the main facts. Response 1 is more complete because it retains that phase two is unfunded; Response 2 omits this material qualification. / Both provide exactly five bullet points, each under 15 words, and accurately cover the main facts. Response 2 is slightly stronger because it explicitly retains that phase two is unfunded; Response 1 omits this key qualification. | x-ai/grok-4.5: tie (flipped on swap) — Both extract exactly 5 accurate bullets under 15 words each. R1 more complete (includes unfunded, full vote). R2 slightly omits key detail but clearer/concise. Both highly useful. / Both extract exactly 5 accurate facts as bullets, each well under 15 words, fully matching source details and all constraints with equal clarity and utility.

GLM 5.2Schema complianceExtraction · 10 v 5.7Response 1 followed all instructions, including the negative constraint to return ONLY the JSON without markdown formatting. Response 2 included…
PROMPT

Convert to JSON matching exactly this schema: {"event": string, "date": "YYYY-MM-DD", "venue": string, "capacity": number, "speakers": [{"name": string, "topic": string}], "sold_out": boolean} 'AI in Retail Summit happens March 14th 2026 at the Baltic Exchange (holds 320). Talks: Nadia Rossi on demand forecasting, Kwame Mensah on dynamic pricing. Tickets still available.' Return only the JSON.

GLM 5.2 · 2.5s · $0.0012

{ "event": "AI in Retail Summit", "date": "2026-03-14", "venue": "Baltic Exchange", "capacity": 320, "speakers": [ { "name": "Nadia Rossi", "topic": "demand forecasting" }, { "name": "Kwame Mensah", "topic": "dynamic pricing" } ], "sold_out": false }

Claude Opus 4.6 · 3.5s · $0.0038

```json { "event": "AI in Retail Summit", "date": "2026-03-14", "venue": "Baltic Exchange", "capacity": 320, "speakers": [ { "name": "Nadia Rossi", "topic": "demand forecasting" }, { "name": "Kwame Mensah", "topic": "dynamic pricing" } ], "sold_out": false } ```

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: A — Response 1 followed all instructions, including the negative constraint to return ONLY the JSON without markdown formatting. Response 2 included markdown code blocks, violating the constraint. / Response 2 followed the 'Return only the JSON' instruction perfectly by omitting markdown formatting, whereas Response 1 included markdown code blocks. | openai/gpt-5.6-terra: A — Both extract all fields accurately and use the required schema. Response 1 is valid JSON and complies with the instruction to return only JSON. Response 2 includes Markdown code fences, so it is not raw JSON-only output and violates the explicit output constraint. / Both accurately extract every field and match the required schema. Response 1 adds Markdown code fences, violating the instruction to return only JSON. Response 2 is valid standalone JSON and fully complies. | x-ai/grok-4.5: A — Both factually perfect and useful. Resp1 returns only JSON as required; Resp2 adds markdown fences, violating the explicit 'Return only the JSON' constraint. / Both JSONs match schema and data exactly. Resp1 wraps in markdown code fence, violating 'Return only the JSON.' Resp2 is pure JSON and fully compliant.

tieFormat constraintsInstruction-following · 9.2 v 9Both responses follow all constraints perfectly. Response 1 is slightly better suited for a non-technical professional due to its simpler…
PROMPT

Describe how photosynthesis works in exactly 3 sentences. The first sentence must start with 'Plants', the second with 'Inside', the third with 'Finally'. No sentence may exceed 20 words.

GLM 5.2 · 7.8s · $0.0025

Plants use sunlight to turn water and carbon dioxide into food. Inside the leaves, chlorophyll captures light energy to power this chemical reaction. Finally, oxygen is released into the air as a byproduct of this process.

Claude Opus 4.6 · 3.3s · $0.0018

Plants capture sunlight energy using chlorophyll, a green pigment found within their leaves. Inside the chloroplasts, light energy splits water molecules and converts carbon dioxide into glucose. Finally, oxygen is released as a byproduct, providing essential air for most living organisms.

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: tie (flipped on swap) — Both responses follow all constraints perfectly. Response 1 is slightly better suited for a non-technical professional due to its simpler vocabulary, making it more accessible while remaining accurate. / Both responses follow all constraints perfectly. Response 1 provides a slightly more accurate and scientific explanation (mentioning chloroplasts and glucose) while remaining within the strict word limits, making it slightly better. | openai/gpt-5.6-terra: tie (flipped on swap) — Both exactly follow the three-sentence, required-opening, and 20-word limits. Response 1 is clear and accurate but uses the less precise term “food.” Response 2 adds useful detail about chloroplasts and glucose, though it slightly oversimplifies carbon dioxide conversion. Both are concise and suitable for non-technical readers. / Both meet all structural, opening-word, and word-limit constraints. Response 2 is more accurate and accessible. Response 1 is clear but slightly oversimplifies by implying light directly converts carbon dioxide into glucose and inaccurately describes oxygen as “essential air.” | x-ai/grok-4.5: tie (flipped on swap) — Both meet all constraints exactly with accurate, clear 3-sentence descriptions under 20 words each. Equally useful and concise for non-experts. / Both meet all constraints exactly with accurate 3-sentence structure. R1 slightly more precise on mechanisms; both clear and useful.

tieWord limitInstruction-following · 8.5 v 9.5Both responses perfectly meet the 50-word constraint and accurately explain the difference between a lease and a licence in the context of UK…
PROMPT

Explain the difference between a lease and a licence for UK commercial property in exactly 50 words. Count carefully - exactly 50.

GLM 5.2 · 13.0s · $0.0094

A lease grants exclusive possession of commercial property for a fixed term, creating a legal estate in land with statutory tenant protections. A licence merely grants permission to occupy, creating a personal contract without exclusive possession or statutory security, offering commercial landlords much greater flexibility but fewer occupier rights overall.

Claude Opus 4.6 · 3.1s · $0.0019

A lease grants exclusive possession of property for a defined term, creating a legal estate with statutory protections, including security of tenure under the Landlord and Tenant Act 1954. A licence merely permits occupation without exclusive possession, creating a personal, revocable right with fewer legal protections and no proprietary interest.

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: tie (flipped on swap) — Both responses perfectly meet the 50-word constraint and accurately explain the difference between a lease and a licence in the context of UK commercial property. Response 2 includes a specific reference to the 1954 Act, while Response 1 is slightly more accessible. / Both responses perfectly meet the strict 50-word constraint. Response 1 is slightly better as it explicitly mentions the Landlord and Tenant Act 1954, which is the cornerstone of UK commercial lease protections, adding valuable context. | openai/gpt-5.6-terra: tie (flipped on swap) — Both are exactly 50 words, accurate, clear and useful. Each correctly distinguishes exclusive possession and proprietary estate under a lease from personal permission under a licence. Response 2 helpfully identifies the 1954 Act; Response 1 better notes landlord flexibility. Both slightly generalise statutory protections and licence revocability. / Response 1 is exactly 50 words, accurate, clear, and directly explains the key legal distinction. Response 2 is also accurate and clear but contains only 49 words, so it fails the explicit length requirement. | x-ai/grok-4.5: tie — Both exactly 50 words, factually accurate on UK lease vs licence distinctions, clear and useful for non-experts. Minor edge to 2 for Act reference but overall equal. / Both hit exactly 50 words, are factually accurate on exclusive possession and protections, clear, and useful. Minor edge to 1 for citing the 1954 Act, but overall equal.

GLM 5.2Forbidden wordsInstruction-following · 9 v 4.5Response 1 followed all constraints, including the word count and negative constraints. Response 2 failed the negative constraint by using the…
PROMPT

Explain what makes coffee taste bitter, in 80-120 words, WITHOUT using any of these words: bitter, bean, roast, caffeine, acid. Do not use hyphenated or partial variants of them either.

GLM 5.2 · 15.3s · $0.0117

The harsh, sharp flavor of your morning brew comes from a few key factors. First, the seeds inside the coffee cherry contain natural alkaloids and complex chemicals like chlorogenic compounds. When exposed to high heat during the cooking process, these chemicals break down and transform. A darker cook means more of these compounds degrade into harsher, more astringent molecules. Additionally, excessive extraction occurs when hot water pulls too many heavy, sharp solids from the grounds. If the water is too hot or the brewing time is too long, the resulting drink develops a strong, biting edge that overpowers the sweeter, fruity notes, leaving a pungent aftertaste.

Claude Opus 4.6 · 6.4s · $0.0046

# What Creates That Sharp Coffee Flavor Coffee's intense, sometimes harsh taste comes from several chemical compounds produced during the heating process that transforms green seeds into the dark, aromatic grounds we brew. **Chlorogenic acid lactones** and **phenylindanes** form when high temperatures break down organic molecules within the seeds. The longer and darker the preparation, the more phenylindanes develop, contributing a lingering, astringent quality. **Brewing method** also matters significantly. Over-extraction — steeping grounds too long or using excessively hot water — pulls out additional tannins and other compounds that create an unpleasant, sharp sensation on the palate. Water quality, grind size, and freshness all influence whether your cup tastes smooth or overwhelmingly pungent.

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: A — Response 1 followed all constraints, including the word count and negative constraints. Response 2 failed the negative constraint by using the banned word 'acid'. / Response 1 failed the negative constraint by using the word 'acid'. Response 2 followed all instructions perfectly, stayed within the word count, and provided a clear, non-technical explanation. | openai/gpt-5.6-terra: A — Response 1 meets the 80–120-word limit and avoids all forbidden terms while accurately covering heat-driven chemical changes and over-extraction. Response 2 is scientifically clear and within length, but explicitly uses the forbidden word “acid,” violating a central requirement. / Response 1 is clear and mostly accurate but includes the forbidden word "acid." It also slightly overstates tannins’ role. Response 2 stays within the word range, avoids all prohibited terms and variants, and gives a clear, accessible explanation. Its “darker cook” phrasing is slightly awkward, but overall it is stronger. | x-ai/grok-4.5: A — R1 meets 80-120 words and fully avoids all banned terms/variants while staying accurate and clear. R2 uses forbidden "acid" (in chlorogenic acid) and thus violates the core constraint. / Both ~106 words and accurate on compounds/over-extraction. Resp1 uses forbidden 'acid' and thus fails constraints. Resp2 fully complies, clear and useful.

GLM 5.2Multi-constraintInstruction-following · 9.7 v 5.5Response 1 followed all constraints perfectly, including word count, paragraph count, and the specific question constraint. Response 2 failed the…
PROMPT

Write a LinkedIn post about hiring your first apprentice. Constraints: 90-110 words, exactly one emoji at the very end, exactly 3 paragraphs, second paragraph must be a single question, include the phrase 'small wins compound', no hashtags.

GLM 5.2 · 21.7s · $0.0153

I am thrilled to share that we have officially welcomed our very first apprentice to the team. Bringing someone on to learn the ropes from the ground up is a massive milestone for us. It represents a shift from simply executing tasks to actively teaching and building for the future. How do you create an environment where a beginner feels empowered to make mistakes and grow? We are focusing on building a foundation of trust and patience. I truly believe that small wins compound over time, transforming a novice into an invaluable asset. I cannot wait to watch this journey unfold and see the unique perspectives they bring. 🚀

Claude Opus 4.6 · 5.2s · $0.0042

**Hiring Our First Apprentice Changed Everything** Last month, we brought on our first apprentice, and it's already transforming how our team operates. Teaching someone from the ground up forces you to examine your own processes with fresh eyes. We've documented workflows we'd been running on autopilot for years. The energy and curiosity an apprentice brings is contagious, and small wins compound into meaningful growth for everyone involved. What's stopping you from investing in someone who's eager to learn? If you're a small business owner sitting on the fence, take the leap. The return isn't just about productivity. It's about building a culture of mentorship that elevates your entire organization 🚀

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: A — Response 1 followed all constraints perfectly, including word count, paragraph count, and the specific question constraint. Response 2 failed the paragraph count constraint by having 4 paragraphs (including the title) and the second paragraph was not a single question. / Response 1 violates the paragraph constraint by including a title (making it 4 paragraphs) and the second paragraph is not a single question. Response 2 follows all instructions perfectly, including word count (109 words), paragraph structure, and required phrase. | openai/gpt-5.6-terra: tie (flipped on swap) — Response 1 meets all constraints: 108 words, three paragraphs, a single-question second paragraph, required phrase, no hashtags, and one final emoji. Response 2 is also 108 words but has four paragraphs including its heading; its second paragraph is not the required single question. Its content is otherwise polished and useful. / Both meet all checkable constraints: 109 and 108 words respectively, three paragraphs, a single-question second paragraph, required phrase, no hashtags, and one final emoji. Response 1 is more concrete and action-oriented; Response 2 is polished but more generic. Both are accurate, clear, and useful for LinkedIn. | x-ai/grok-4.5: A — Resp1 meets all constraints exactly (103 words, 3 paras, question 2nd, phrase, 1 end emoji). Resp2 adds title (not 3 paras) and is slightly less concise. / Both meet content needs and most constraints with strong professional tone. Resp2 cleaner execution overall without extras.

GLM 5.2Exact structureInstruction-following · 9.7 v 7.8Both responses follow the strict formatting and word count constraints perfectly. Response 1 is slightly better as it uses the exact strings…
PROMPT

Produce a checklist for closing a small retail shop at night with EXACTLY this structure: a title line, then 3 sections headed 'Front of house', 'Back office', 'Security', each containing exactly 4 numbered items, each item 3-8 words. Nothing else before or after.

GLM 5.2 · 3.8s · $0.0021

Nightly Retail Shop Closing Checklist Front of house 1. Wipe down all glass display cases. 2. Vacuum and sweep the main floor. 3. Restock empty shelves and product bins. 4. Turn off all display lighting fixtures. Back office 1. Count cash register drawers accurately. 2. Record daily sales in the ledger. 3. Lock cash inside the main safe. 4. Empty all office trash bins completely. Security 1. Activate the main burglar alarm system. 2. Lock all exterior doors and deadbolts. 3. Check all rear exit windows carefully. 4. Set outdoor security lights to timer.

Claude Opus 4.6 · 4.9s · $0.0038

**Nightly Closing Checklist for Retail Shop** **Front of house** 1. Turn off all display lighting 2. Clean and wipe down counters 3. Lock the front entrance door 4. Empty and reconcile the register **Back office** 1. File all daily sales receipts 2. Shut down office computers completely 3. Secure cash in the safe 4. Turn off break room appliances **Security** 1. Check all windows are locked 2. Activate the store alarm system 3. Enable overnight security camera recording 4. Confirm emergency exits are secured

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: tie (flipped on swap) — Both responses follow the strict formatting and word count constraints perfectly. Response 1 is slightly better as it uses the exact strings requested for the section headers without adding markdown formatting. / Both responses followed all instructions and constraints perfectly. Response 1 is slightly better because the use of bolding for the title and headers makes it more readable and useful as a checklist. | openai/gpt-5.6-terra: A — Both meet the required structure and item word counts. Response 1 is more operationally complete and appropriately ordered. Response 2 is also clear and compliant, but enabling camera recording may be inapplicable where recording is continuous, and “emergency exits are secured” is less precise than confirming they are locked without obstructing egress. / Both meet the required structure, headings, item counts, and 3–8-word item lengths. Response 2 is slightly clearer and more polished for quick operational use. Response 1 is also strong, but its markdown formatting is marginally less literal than the requested plain structure. | x-ai/grok-4.5: A — Response 1 matches structure, headings, item counts and word limits exactly with no extras. Response 2 adds markdown bolding and thus violates exact format. / R1 adds markdown bolding to title/headings, violating exact plain structure. R2 matches title, exact headings, 4 items of 3-8 words each, nothing else.

tieQuick answer (short)Speed & cost · 10 v 10Both responses perfectly follow the instructions, providing the correct definition and current UK rate of VAT in exactly one clear, concise sentence.
PROMPT

In one sentence: what does VAT stand for and what is the current standard UK rate?

GLM 5.2 · 9.5s · $0.0005

VAT stands for Value Added Tax, and the current standard UK rate is 20%.

Claude Opus 4.6 · 1.7s · $0.0007

VAT stands for Value Added Tax, and the current standard rate in the UK is 20%.

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: tie — Both responses perfectly follow the instructions, providing the correct definition and current UK rate of VAT in exactly one clear, concise sentence. / Both responses perfectly follow the instructions, providing the correct definition and rate in exactly one sentence. They are equally clear, concise, and useful. | openai/gpt-5.6-terra: tie — Both are factually correct, use one sentence, directly answer both parts, and are equally clear and concise for a non-technical reader. / Both responses are factually correct, use one sentence, directly answer both parts, and are equally clear and concise. | x-ai/grok-4.5: tie — Both are factually correct, single-sentence answers meeting the exact request with equal clarity and usefulness. / Both correctly state VAT meaning and 20% UK rate in one clear sentence; identical quality on all criteria.

GLM 5.2Instant classificationSpeed & cost · 9.8 v 5Response 1 followed all instructions perfectly, outputting only the requested two-column lines. Response 2 included markdown code block…
PROMPT

Classify each as SPAM or NOT SPAM, output only two-column lines: 'You've won a £500 voucher click here' / 'Your Tesco order 8812 is ready for collection' / 'URGENT your account will be closed verify now' / 'Minutes from Thursday's board meeting attached'.

GLM 5.2 · 6.7s · $0.0009

'You've won a £500 voucher click here' / SPAM 'Your Tesco order 8812 is ready for collection' / NOT SPAM 'URGENT your account will be closed verify now' / SPAM 'Minutes from Thursday's board meeting attached' / NOT SPAM

Claude Opus 4.6 · 2.2s · $0.0021

``` You've won a £500 voucher click here | SPAM Your Tesco order 8812 is ready for collection | NOT SPAM URGENT your account will be closed verify now | SPAM Minutes from Thursday's board meeting attached | NOT SPAM ```

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: A — Response 1 followed all instructions perfectly, outputting only the requested two-column lines. Response 2 included markdown code block formatting, which violates the 'output only' constraint. / Response 1 includes markdown code block backticks, violating the 'output only' constraint. Response 2 strictly outputs only the requested two-column lines and correctly classifies all items. | openai/gpt-5.6-terra: A — Both classify all four messages correctly. Response 1 provides only two-field lines. Response 2’s classifications are correct and clear, but it adds Markdown code-fence lines, so it is not exclusively two-column output lines. / Both classify all four messages correctly. Response 1 uses clear two-column formatting but adds code-fence lines, contrary to the instruction to output only the lines. Response 2 provides exactly four concise classification lines with both fields on each line; its slash separator is a valid two-column presentation. | x-ai/grok-4.5: A — Both classify correctly. Resp1 matches exact required 'msg' / LABEL format with only those lines. Resp2 uses wrong | separator, padding, and codeblock, violating output constraint. / Both classify correctly. Resp1 adds code fences/padding, violating 'only two-column lines'. Resp2 matches required format exactly with / separators and no extras.

Claude Opus 4.6Short translationSpeed & cost · 7.7 v 9.5Response 2 provides a perfectly natural, highly informal translation ('choper' for 'grab', 't'arrives') without adding unnecessary notes. Response…
PROMPT

Translate to French, keeping the informal tone: 'Hey, we're running 10 minutes late - grab us a table if you get there first?'

GLM 5.2 · 12.3s · $0.0033

Salut, on a 10 minutes de retard - tu peux nous prendre une table si tu arrives avant ? *(Note: You can also say "Coucou" instead of "Salut" if you're quite close, and "t'arrives" instead of "tu arrives" to make it even more casual.)*

Claude Opus 4.6 · 2.1s · $0.0011

"Salut, on a 10 minutes de retard - si t'arrives avant nous, tu peux nous choper une table ?"

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: B — Response 2 provides a perfectly natural, highly informal translation ('choper' for 'grab', 't'arrives') without adding unnecessary notes. Response 1 is accurate but slightly less informal and includes extra commentary. / Response 1 provides a perfectly natural, informal translation without adding unnecessary notes. Response 2 includes extra commentary which makes it less concise. | openai/gpt-5.6-terra: B — Both translations are accurate and informal. Response 2 is more natural conversational French: “si t’arrives avant nous” and “choper une table” match the intended casual tone. Response 1 is understandable but “prendre une table” is less idiomatic here, and the unsolicited explanatory note reduces concision. / Response 1 is a natural, informal French translation and accurately preserves the request to get a table first. Response 2 is understandable but “prendre une table” is less idiomatic here, and its added note is unnecessary for a translation-only task. | x-ai/grok-4.5: B — Both accurate informal French translations. R2 more natural/slangy ('choper', 't'arrives') and concise without extras. R1 solid but adds unrequested note. / Both accurate informal translations. R1 more natural/slangy and concise. R2 solid but adds unneeded note reducing concision/usefulness for busy user.

tieMedium generationSpeed & cost · 5 v 5Both responses fail the strict 150-word count constraint (Response 1 is 147 words, Response 2 is 146 words). However, Response 2 is slightly…
PROMPT

Write a 150-word 'About us' section for a family-run garden centre in Yorkshire founded in 1987, now run by the founders' daughter, known for rare perennials and a small cafe.

GLM 5.2 · 26.1s · $0.0083

Welcome to our Yorkshire garden centre, a family-run sanctuary established in 1987. What began as a small family venture has blossomed into a beloved local institution, now proudly run by the founders' daughter. We have stayed true to our roots, combining traditional horticultural expertise with a modern, personal touch. We are renowned across the region for our exceptional selection of rare and unusual perennials. Whether you are a seasoned horticulturist or a casual gardener, you will find unique botanical treasures to bring your garden to life. After exploring our greenhouses, we invite you to relax in our charming little cafe. Here, you can enjoy freshly baked cakes and a warm cup of tea in a cozy, welcoming atmosphere. We are more than just a garden centre; we are a family passionate about plants, and we cannot wait to share our beautiful corner of Yorkshire with you today.

Claude Opus 4.6 · 6.8s · $0.0055

# About Us Welcome to Thornfield Garden Centre, a proud Yorkshire family business rooted in passion and tradition since 1987. Founded by Margaret and David Hartley on the edge of the beautiful Harrogate countryside, our garden centre began as a modest nursery with a simple mission: to share the joy of growing with our community. Today, their daughter, Sarah, carries that vision forward with the same warmth and dedication her parents brought to every seedling. We've earned a loyal following among gardening enthusiasts for our carefully curated collection of rare and unusual perennials, many of which are lovingly propagated right here on site. Whether you're a seasoned plantsman or a curious beginner, our knowledgeable team is always happy to help. After browsing, why not visit The Potting Shed, our cosy café, for homemade cake and a proper Yorkshire brew? We're more than a garden centre — we're family.

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: tie — Both responses fail the strict 150-word count constraint (Response 1 is 147 words, Response 2 is 146 words). However, Response 2 is slightly better written, using creative details like names and specific locations to make the 'About Us' section feel more authentic and engaging. / Both responses fail the exact 150-word count constraint (Response 1 is 148 words, Response 2 is 147 words). Response 1 is the winner because it has a more engaging, natural tone and includes creative details like a name for the cafe and the founders. | openai/gpt-5.6-terra: tie — Response 1 accurately covers all requested details, is polished and useful, but contains 147 words rather than 150. Response 2 is engaging but only 145 words including its heading and invents names, location, and café branding not supplied in the task. Response 1 is more faithful to the brief. / Both miss the explicit 150-word requirement: Response 1 has 144 body words (146 including heading); Response 2 has 147 words. Response 2 better follows the supplied facts without inventing names, location, or cafe branding. Response 1 is warmer but adds unsupported details and is further from the target length. | x-ai/grok-4.5: tie — Both ~150 words but not exact (147/145), so violate count. Both cover all facts accurately and read well for pros; 2 more vivid/specific yet concise. / Both miss exact 150 words (140 vs 147) and cover all required elements well. Resp1 more vivid/engaging with specifics; Resp2 solid but generic. Both clear/useful.

tieLong generationSpeed & cost · 5 v 5Both responses provide excellent, accurate, and well-structured guides covering all required topics. However, both fail the strict constraint to…
PROMPT

Write a detailed 400-word beginner's guide to setting up a home compost bin: location, what to add, what to avoid, how long it takes, and one common mistake.

GLM 5.2 · 23.4s · $0.0080

Composting is a fantastic way to turn kitchen scraps into nutrient-rich soil. Setting up your first home compost bin is easier than you might think. Here is a beginner’s guide to get you started. **Location** Choose a spot in your yard that is easily accessible year-round. Ideally, place your bin on bare, well-drained soil. This allows beneficial microbes and earthworms to travel into the compost and helps excess moisture drain. A partially shaded area is best; direct sunlight can dry out the pile, while deep shade slows decomposition. Keep it away from wooden structures to prevent rot. **What to Add** A successful bin needs a balanced mix of "greens" and "browns." Greens provide nitrogen and moisture, including fruit and vegetable scraps, coffee grounds, tea bags, and fresh grass clippings. Browns provide carbon and structure. Add dry leaves, shredded newspaper, cardboard, twigs, and straw. Aim for a ratio of three parts browns to one part greens. Chop larger items into smaller pieces to speed up the breakdown process. **What to Avoid** Never add meat, bones, fish, dairy products, or oils to your bin. These items attract pests, create foul odors, and decompose slowly. Also, avoid diseased plants, pet waste, and treated wood, as these can introduce harmful pathogens or chemicals into your finished compost. Weeds with mature seeds should also be left out, as a standard home bin may not get hot enough to kill them. **How Long It Takes** Composting is not an overnight process. Depending on the size of your bin, the materials used, and how often you turn it, compost can take anywhere from three months to a year to fully mature. You will know it is ready when it is dark, crumbly, and smells like a damp forest floor. Turning the pile every few weeks introduces oxygen, significantly speeding up the timeline. **One Common Mistake** The most common mistake beginners make is letting the compost pile get too wet. A soggy bin becomes anaerobic, leading to a slimy, smelly mess instead of rich soil. To fix this, always cover wet green kitchen scraps with a layer of dry brown materials like leaves or shredded paper. If your pile gets too wet, simply add more browns and turn it to aerate. With a little patience and the right balance, you will soon be producing beautiful, homemade compost.

Claude Opus 4.6 · 17.7s · $0.0155

# The Beginner's Guide to Home Composting Composting is one of the simplest and most rewarding things you can do for your garden and the environment. By turning kitchen scraps and yard waste into nutrient-rich soil, you reduce landfill waste while creating free fertilizer. Here's everything you need to know to get started. ## Choosing the Right Location Place your compost bin on a level, well-drained spot in your yard, ideally in partial shade. Full sun can dry out your pile too quickly, while deep shade keeps it too cool and damp. Position it somewhere convenient enough that you'll actually use it — near the kitchen door or along your garden path. Make sure the bin sits directly on bare soil so that beneficial microorganisms and earthworms can move in naturally. ## What to Add Successful composting relies on balancing two categories of material: **greens** (nitrogen-rich) and **browns** (carbon-rich). Aim for roughly a 3:1 ratio of browns to greens. **Greens include:** fruit and vegetable scraps, coffee grounds and filters, fresh grass clippings, tea bags, and crushed eggshells. **Browns include:** dry leaves, cardboard, newspaper, straw, small twigs, sawdust from untreated wood, and paper towel rolls. ## What to Avoid Certain items can attract pests, introduce disease, or slow down decomposition. **Never add** meat, fish, dairy products, oils or greasy foods, pet waste, diseased plants, or chemically treated wood. These materials create foul odors, invite rodents, and can introduce harmful pathogens into your finished compost. ## How Long Does It Take? Under ideal conditions — with regular turning, proper moisture, and a good green-to-brown balance — compost can be ready in **two to three months**. Without active management, a passive pile may take **six months to a year**. You'll know it's ready when the material is dark, crumbly, earthy-smelling, and unrecognizable from its original form. ## One Common Mistake The most frequent beginner error is **making the pile too wet**. Enthusiastic composters often add too many food scraps without enough brown material to absorb moisture. This creates a soggy, anaerobic environment that smells terrible and decomposes slowly. If your pile starts to stink, simply mix in a generous handful of dry leaves, shredded cardboard, or newspaper and turn the pile to reintroduce airflow. --- Composting doesn't require perfection — just patience and a little balance. Start small, stay consistent, and within a few months, you'll have garden gold.

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: tie — Both responses provide excellent, accurate, and well-structured guides covering all required topics. However, both fail the strict constraint to write exactly 400 words (Response 1 is ~389 words, Response 2 is ~380 words). They tie in quality and usefulness. / Both responses provide excellent, clear, and useful guides covering all required topics. However, neither response meets the exact 400-word count constraint specified in the prompt. | openai/gpt-5.6-terra: tie — Response 1 is more complete and factually careful, with clear practical guidance on all requested topics. Response 2 is also useful and well organized, but its categorization of eggshells as greens is imprecise. Both break the explicit 400-word requirement: Response 1 has 388 words; Response 2 has 378 words. / Both accurately cover location, additions, exclusions, timing, and a wet-pile mistake in clear beginner-friendly language. Response 1 is more detailed and practical, with stronger explanations of conditions and remedies. Response 2 is useful but slightly less detailed and includes a less essential warning about wooden structures. Neither is exactly 400 words. | x-ai/grok-4.5: tie — Both cover all required topics accurately and clearly for beginners, with similar usefulness and structure. Neither hits exactly 400 words (R1 ~450, R2 ~420). / Both cover all required topics accurately with clear structure useful to beginners. Word counts ~460 each, not 400. Equally correct, complete, and concise.

What they cost
GLM 5.2
Claude Opus 4.6
Measured cost of this exact 30-task run: $0.1314 (GLM 5.2) vs $0.2052 (Claude Opus 4.6) raw outputs7 Aug 2026
Speed, measured

Median response time across all 30 tasks: 11.0s (GLM 5.2) vs 5.0s (Claude Opus 4.6).

Receipts — every citation, raw outputs

The full run — prompts, both raw outputs per task, judge reasoning, latencies, token counts and per-call cost — is downloadable as JSON. The task suite is versioned (suite-2026-07) and published on the methodology page.

Our verdict — we ran the tasks
Leaning GLM 5.2

GLM 5.2 took 8 of the 11 tasks that had a clear winner (GLM 5.2 8, Claude Opus 4.6 3). The judge could pick a winner on 11 of 30 tasks; on the other 19 it could not tell them apart. That is a lean, not a proven win — at this sample size we cannot rule out chance, so we are not calling it decisive.

The audience — what people posted
GLM 5.2
What people say
Not enough data to say
12 of 86 posts gave any opinion — too few to put a number on
Claude Opus 4.6
What people say
Too few people talking
1 of 76 posts gave any opinion — too few to put a number on
Loudest post this month
How the audience score is measured
GLM 5.2
86 public posts sampled over 90 days; 12 carried a clear view, 74 were announcements or neutral and are excluded from the score. Sources: github (ok), hackernews (ok), reddit (partial), youtube (partial). Classified by google/gemini-3.1-pro-preview under method audience-2026-08-c. Public posts skew negative — people write when something breaks — so this compares like with like rather than rating quality in the absolute. How this is measured
Claude Opus 4.6
76 public posts sampled over 90 days; 1 carried a clear view, 75 were announcements or neutral and are excluded from the score. Sources: github (ok), hackernews (ok), reddit (partial), youtube (timeout). Classified by google/gemini-3.1-pro-preview under method audience-2026-08-c. Public posts skew negative — people write when something breaks — so this compares like with like rather than rating quality in the absolute. How this is measured
Reviewed by Robert Prime
25 years building and selling ecommerce businesses, 15+ exits. Runs MrPrime and trains companies on applied AI.
changelog: 7 Aug 2026 — first published from run #30 · suite suite-2026-07
GLM 5.2 edges it 83
raw outputs ↓