Battles / Best-value API

Kimi K3 wins 11–3

a solid win on the tasks that separated them (14 of 30 tasks were decisive) — close enough that the loser is still worth a look.

Kimi K3 vs Claude Opus 4.6 · Best-value APImodels in this battle: Kimi K3 vs Claude Opus 4.6

tested 10 Aug 202630 tasks3 judges from different labs
Show the full workings
last verified 10 Aug 2026suite suite-2026-07judges: google/gemini-3.1-pro-preview, openai/gpt-5.6-terra, x-ai/grok-4.5 (none of them a contestant)judge protocol judge-2026-08 — constraint cap enforced in code, applied to 71 judge scores in this runnot statistically decisive — Wilcoxon signed-rank on score margins p=0.0128 (n=25); sign test on win counts p=0.0574 — held to our confidence gatejudge swap-consistency (Cohen’s κ) = 0.579moderate agreement κ measures how often a judge reached the same verdict when the two answers were shown in the opposite order, corrected for agreement you would get by chance. A battle below 0.21 does not publish a winner at all; above it, this number is how much of the scoreline survived the swap test rather than a claim that the judges were reliable.

The evidence

Suite-by-suite
Writing
221
Coding
140
Reasoning
230
Extraction
221
Instruction-following
320
Speed & cost
131

blue = Kimi K3 wins · grey = ties · white = Claude Opus 4.6 wins (5 tasks per suite)

Round-by-round — all 30 tasks
Kimi K3Cold emailWriting · 9.7 v 8.8Both responses follow all instructions and constraints perfectly, staying under the 120-word limit. Response 1 is slightly better because it…
PROMPT

Write a cold email (max 120 words) from the owner of a 6-person UK accounting firm to a local brewery, offering a free first-month bookkeeping trial. Friendly but professional, one clear call to action, no buzzwords.

Kimi K3 · 69.3s · $0.0378

**Subject:** A free month of bookkeeping for [Brewery name] Hi [Name], I'm Sarah, owner of [Firm name] — a six-person accounting practice just down the road in [Town]. We look after a few local food and drink businesses, so we know the paperwork breweries deal with: beer duty returns, VAT on taproom and wholesale sales, and stock that never sits still. I'd like to offer you a free first month of bookkeeping. We'll bring your records up to date, reconcile the bank, and hand you clean numbers — no charge, no obligation. Fancy a quick chat? Reply to this email and we'll set up a 15-minute call at a time that suits you. Best wishes, Sarah [Surname] [Firm name] | [Phone]

Claude Opus 4.6 · 6.3s · $0.0046

**Subject: Happy to help with your books – first month on us** Hi, I'm Sarah, owner of Clearpath Accounting here in [town]. We're a small team of six, and we work with several food and drink businesses locally. I know running a brewery means your time goes into making great beer, not reconciling invoices. We'd like to offer you a free month of bookkeeping – no strings, no contract. It's simply how we prefer to introduce ourselves. If the fit works, we can talk about continuing. If not, you walk away with a tidy set of books. Would you be free for a 15-minute call this week or next? Best, Sarah Mitchell Clearpath Accounting [phone] | [email]

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: A — Both responses follow all instructions and constraints perfectly, staying under the 120-word limit. Response 1 is slightly better because it includes specific, highly relevant details for a UK brewery (beer duty returns, VAT), making the cold email more compelling. / Both responses follow all instructions perfectly and stay under the 120-word limit. Response 2 wins because it includes specific, highly relevant details for a brewery (beer duty, taproom VAT), making it a much more effective cold email. | openai/gpt-5.6-terra: A — Both meet the 120-word limit and all core requirements. Response 1 is more specifically tailored to a UK brewery, with credible bookkeeping examples and a direct single CTA. Response 2 is friendly, clear, and compliant but more generic and slightly less locally tailored. / Both meet the length, tone, sender, local-brewery, free-trial, and single-CTA requirements. Response 2 is more tailored to brewery bookkeeping and states the trial scope clearly, making it more useful. Response 1 is strong but slightly more generic. | x-ai/grok-4.5: tie (flipped on swap) — Both under 120 words, meet all style rules with one clear CTA. Resp1 stronger on brewery-specific relevance and professionalism; Resp2 solid but slightly generic. / Both under 120 words, meet all constraints (friendly/pro, one CTA, no buzzwords, free trial offer from 6-person UK firm). Equally correct, clear, useful and concise.

Kimi K3Product descriptionWriting · 9.2 v 8.5Both responses follow all instructions and constraints perfectly, including the word count. Response 1 is slightly more engaging and natural in…
PROMPT

Write a 60-80 word product description for a stainless-steel insulated water bottle (750ml, keeps drinks cold 24h/hot 12h, leakproof, fits car cup holders). Target audience: gym-goers. Avoid cliches like 'stay hydrated in style'.

Kimi K3 · 21.5s · $0.0690

Your last set deserves water as cold as your first. This 750ml stainless-steel bottle keeps drinks ice-cold for 24 hours and coffee hot for 12—pre-workout included. The leakproof lid survives being tossed in your gym bag next to dry clothes. Double-wall insulation means zero condensation, so your grip stays secure mid-set. It fits standard car cup holders, and the steel body shrugs off drops, dents, and daily abuse.

Claude Opus 4.6 · 4.9s · $0.0033

# HydroCore 750ml Insulated Water Bottle Engineered for athletes who demand more from their gear. This double-wall stainless-steel bottle locks in ice-cold temperatures for 24 hours or keeps your pre-workout hot for 12. The leakproof seal means it tosses into any gym bag without worry, while the slim profile slides right into car cup holders for easy transport between sessions. Durable, sweat-free, and built to survive every rep, set, and sprint.

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: A — Both responses follow all instructions and constraints perfectly, including the word count. Response 1 is slightly more engaging and natural in its tone for the target audience, making it the winner. / Both responses follow all instructions and constraints perfectly. Response 2 wins for having a slightly more engaging and punchy opening hook tailored specifically to the gym-goer audience. | openai/gpt-5.6-terra: tie (flipped on swap) — Both meet the 60–80 word range, include all core specifications, target gym-goers, and avoid the banned cliché. Response 2 is more polished and concise, with clearer benefit-led phrasing. Response 1 adds unsupported claims about drop/dent resistance and has a slightly awkward “next to dry clothes” line. / Both meet the 60–80-word requirement, target gym-goers, include all specified features, and avoid the prohibited cliché. Response 2 is more immediately engaging and benefit-led for the audience. Both add unverified durability/condensation claims; Response 1 also has slightly less natural phrasing around hot pre-workout. | x-ai/grok-4.5: A — Both hit 60-80 words, cover all specs, target gym users, avoid banned cliche. R1 more vivid and gym-specific without fluff; R2 solid but generic title and slightly less punchy. / Both hit 60-80 words, cover all specs, target gym users, avoid cliches. R2 more vivid and useful; R1 solid but slightly less punchy.

Claude Opus 4.6Summarise messy notesWriting · 7.5 v 9.3Response 2 strictly adheres to the 5-bullet format by incorporating all information into the bullets, whereas Response 1 adds an extra line…
PROMPT

Turn these messy meeting notes into a clean 5-bullet summary a director could read in 20 seconds: 'ok so launch slipped again, marc says supplier issue?? maybe 3 wks. sarah wants to push the email campaign anyway but tom thinks wait. budget - we're 8k over, mostly the packaging redo. Q: do we tell retail partners now or after new date confirmed. also NEED to hire the warehouse temp before august rush. next mtg tues.'

Kimi K3 · 26.6s · $0.0116

- **Launch slipped again** — ~3-week delay due to supplier issue (per Marc); new date TBD - **Email campaign: decision needed** — Sarah wants to proceed as planned; Tom wants to wait for new launch date - **Budget: $8K over** — primarily from packaging redo - **Retail partner comms: open question** — notify now or wait until new date is confirmed - **Action: hire warehouse temp** — must be in place before August rush *Next meeting: Tuesday*

Claude Opus 4.6 · 5.5s · $0.0037

• **Launch delayed ~3 weeks** due to a supplier issue (Marc investigating); revised date TBC. • **Marketing timing split:** Sarah proposes running the email campaign now; Tom recommends holding until the new date is locked — decision needed. • **Budget is $8K over**, driven primarily by the packaging redesign. • **Retail partner communication** pending — team needs to decide whether to notify now or wait until the new launch date is confirmed. • **Immediate action items:** Hire warehouse temp before the August rush; next meeting Tuesday to resolve open decisions.

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: B — Response 2 strictly adheres to the 5-bullet format by incorporating all information into the bullets, whereas Response 1 adds an extra line outside the bullets. Both are highly accurate, clear, and useful. / Response 1 perfectly integrates all information into exactly five bullets. Response 2 includes an extra line outside the five bullets, making it slightly less compliant with the '5-bullet summary' format, though both are highly readable and accurate. | openai/gpt-5.6-terra: B — Both are accurate, concise, and provide five clear bullets. Response 1 adds the meeting date outside the bullets, while Response 2 incorporates all key information within exactly five bullets. Response 2 includes a minor unsupported detail (“Marc investigating”), but is slightly more useful because it consolidates next steps and timing. / Response 1 is an accurate, director-ready five-bullet summary and includes all key decisions, risks, actions, and meeting timing. Response 2 is also accurate and clear, but places the next-meeting detail outside the five bullets, so it does not strictly deliver the requested five-bullet-only format. | x-ai/grok-4.5: tie (flipped on swap) — Both accurately condense notes into exactly 5 clear bullets useful for a director. R1 slightly cleaner by isolating next meeting; R2 folds it in acceptably. No constraints broken. / Both accurate and useful. R1 is exactly 5 clean bullets including all points. R2 adds extra line beyond 5 bullets, slightly less concise.

tieTone rewriteWriting · 9.5 v 9.5Both responses perfectly follow the instructions, maintaining all original facts while adopting a warm, accountable tone. Both are well under the…
PROMPT

Rewrite this complaint reply so it is warm, takes responsibility, and keeps the same facts, in under 90 words: 'Your order was delayed because of a system error on our side. The refund has been processed and will arrive in 3-5 days. We cannot offer further compensation. Let us know if you have questions.'

Kimi K3 · 28.6s · $0.0115

We're truly sorry your order was delayed — this was caused by a system error on our end, and we take full responsibility for the inconvenience. Your refund has been processed and should arrive within 3-5 days. While we're unable to offer additional compensation, we completely understand your frustration and hope you'll give us another chance to serve you better. If you have any questions at all, please don't hesitate to reach out — we're here to help.

Claude Opus 4.6 · 4.1s · $0.0030

We're truly sorry for the delay with your order — this was caused by a system error on our end, and we take full responsibility for the inconvenience. We've processed your refund, which should arrive within 3–5 days. While we're unable to offer additional compensation at this time, please know we value your patience and your business. Don't hesitate to reach out if you have any questions — we're here to help and want to make sure you're taken care of.

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: tie — Both responses perfectly follow the instructions, maintaining all original facts while adopting a warm, accountable tone. Both are well under the 90-word limit and highly useful for a professional setting. / Both responses perfectly follow all instructions, keeping the original facts while adopting a warm, accountable tone. Both are well under the 90-word limit and are highly useful and clear. | openai/gpt-5.6-terra: tie (flipped on swap) — Both preserve the key facts, accept responsibility, use a warm tone, and stay under 90 words. Response 2 is slightly more polished and customer-focused, with smoother phrasing and reassurance. Response 1 is also strong but adds "hope you'll give us another chance," which is somewhat less directly useful. / Both are warm, responsible, accurate, and under 90 words. Response 2 is slightly more concise and preserves the compensation limitation more faithfully. Response 1’s phrase “at this time” could imply compensation may be available later, slightly softening the original position. | x-ai/grok-4.5: tie — Both warm, responsible, fact-preserving, clear, under 90 words, equally useful. / Both rewrites are warm, own the error, preserve all facts exactly, and stay well under 90 words with clear professional tone.

tieHeadlinesWriting · 9 v 9.5Both models followed all instructions and constraints perfectly. Response 1 is slightly better because its curiosity-driven headlines still…
PROMPT

Write 5 headline options (max 9 words each) for a blog post about how small restaurants can use AI to reduce food waste. Mix: 2 practical, 2 curiosity-driven, 1 with a number.

Kimi K3 · 44.4s · $0.0252

**Practical** 1. How Small Restaurants Can Cut Food Waste with AI 2. A Small Restaurant's Guide to AI-Powered Waste Reduction **Curiosity-Driven** 3. Why Smart Restaurants Are Letting AI Watch Their Trash 4. The AI Trick Cutting Restaurant Food Waste in Half **Number** 5. 5 Ways AI Helps Small Restaurants Slash Food Waste All options are 9 words or fewer. Want me to adjust the tone (more formal, punchier, etc.) or generate alternates for any category?

Claude Opus 4.6 · 4.9s · $0.0033

# 5 Headline Options ## Practical 1. **How Small Restaurants Use AI to Slash Food Waste** 2. **AI Tools That Help Small Kitchens Waste Less Food** ## Curiosity-Driven 3. **Your Dumpster Knows Something Your Menu Doesn't** 4. **The Surprising Way Tiny Restaurants Are Outsmarting Waste** ## Numbered 5. **3 AI Tricks Saving Small Restaurants From Wasted Food**

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: tie (flipped on swap) — Both models followed all instructions and constraints perfectly. Response 1 is slightly better because its curiosity-driven headlines still clearly communicate the core topic (AI), whereas Response 2's curiosity headlines omit the mention of AI entirely. / Both responses follow all instructions and constraints perfectly. Response 1 is slightly better because its curiosity-driven headlines are more engaging, and it avoids the unnecessary conversational filler included at the end of Response 2. | openai/gpt-5.6-terra: B — Both provide exactly five headlines, with the requested mix and all headlines within nine words. Response 2 is cleaner and more concise, with stronger varied hooks and no unsupported performance claim. Response 1 is useful, but “cutting food waste in half” is a potentially unsubstantiated factual claim. / Both provide exactly five headlines, with two practical, two curiosity-driven, and one numbered option; all headlines meet the nine-word limit. Response 1 is slightly stronger: its categories are cleaner and its curiosity headlines are more distinctive. Response 2 adds unnecessary follow-up text and one curiosity headline reads more like a performance claim. | x-ai/grok-4.5: abstain — unparseable

tieBug fixCoding · 9 v 9.7Both responses perfectly fix the bugs (including the implicit numeric sort bug), avoid mutating the input, and follow the constraint to provide…
PROMPT

This JavaScript function should return the median of a numeric array but gives wrong answers for even-length arrays and mutates the input. Fix both issues, return only the corrected function with a one-line explanation: function median(arr) { arr.sort(); const mid = Math.floor(arr.length / 2); return arr[mid]; }

Kimi K3 · 4.4s · $0.0087

```javascript function median(arr) { const sorted = [...arr].sort((a, b) => a - b); const mid = Math.floor(sorted.length / 2); return sorted.length % 2 ? sorted[mid] : (sorted[mid - 1] + sorted[mid]) / 2; } ``` **Explanation:** Sorts a copy of the array (with a numeric comparator) to avoid mutating the input, and averages the two middle elements when the length is even.

Claude Opus 4.6 · 3.2s · $0.0034

```js function median(arr) { const sorted = [...arr].sort((a, b) => a - b); const mid = Math.floor(sorted.length / 2); return sorted.length % 2 !== 0 ? sorted[mid] : (sorted[mid - 1] + sorted[mid]) / 2; } ``` Copies the array to avoid mutation, uses a numeric comparator for correct sorting, and averages the two middle elements for even-length arrays.

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: tie — Both responses perfectly fix the bugs (including the implicit numeric sort bug), avoid mutating the input, and follow the constraint to provide only the function and a one-line explanation. / Both responses perfectly fix the issues (including the implicit numeric sort bug), avoid mutating the input, handle even-length arrays correctly, and follow the formatting constraint exactly. | openai/gpt-5.6-terra: tie — Both correctly sort a copy using a numeric comparator, preserve the input, and average the two middle values for even-length arrays. Each provides only the corrected function plus a one-line explanation. Response 2 is marginally more concise, but both are equally correct and useful. / Both correctly avoid mutation, use numeric sorting, and average middle values for even-length arrays. Each provides only the corrected function plus a one-line explanation, making them clear and useful. | x-ai/grok-4.5: tie (flipped on swap) — Both correctly fix mutation via copy+numeric sort and even-length avg. Resp2 better matches 'only fn + one-line expl' exactly; Resp1 adds bold label/extra formatting. / Both correctly fix mutation via copy, numeric sort, and even-length average. Both supply function plus one-line explanation as required. Nearly identical and equally clear/useful.

tieSmall featureCoding · 9.5 v 9.3Both responses correctly implement the requested logic and include the required test cases. Response 2 provides a more elegant and concise…
PROMPT

Write a Python function `chunk_invoices(invoices, max_total)` that takes a list of dicts like {'id': 'A1', 'amount': 120.5} and groups them into batches where each batch's summed amount does not exceed max_total. A single invoice larger than max_total goes in its own batch. Preserve order. Include 3 test cases using assert.

Kimi K3 · 74.7s · $0.0453

```python def chunk_invoices(invoices, max_total): """Group invoices into ordered batches whose summed amounts <= max_total. An invoice individually larger than max_total is placed in its own batch. """ batches = [] current_batch = [] current_total = 0 for invoice in invoices: amount = invoice['amount'] # Oversized invoice: flush current batch, then isolate it if amount > max_total: if current_batch: batches.append(current_batch) current_batch, current_total = [], 0 batches.append([invoice]) continue # Would exceed the limit -> close this batch and start a new one if current_total + amount > max_total: batches.append(current_batch) current_batch, current_total = [], 0 current_batch.append(invoice) current_total += amount if current_batch: batches.append(current_batch) return batches # Test 1: greedy filling; an exact-fit amount closes the batch inv = [ {'id': 'A1', 'amount': 100}, {'id': 'A2', 'amount': 100}, {'id': 'A3', 'amount': 100}, ] assert chunk_invoices(inv, 200) == [[inv[0], inv[1]], [inv[2]]] # Test 2: invoice larger than max_total gets its own batch; order preserved inv = [ {'id': 'B1', 'amount': 50}, {'id': 'B2', 'amount': 500}, {'id': 'B3', 'amount': 40}, ] assert chunk_invoices(inv, 100) == [[inv[0]], [inv[1]], [inv[2]]] # Test 3: float amounts, mid-list overflow, and exact-boundary fit inv = [ {'id': 'C1', 'amount': 120.5}, {'id': 'C2', 'amount': 60.25}, {'id': 'C3', 'amount': 30.0}, {'id': 'C4', 'amount': 170.0}, ] result = chunk_invoices(inv, 200.0) assert result == [[inv[0], inv[1]], [inv[2], inv[3]]] assert all(sum(i['amount'] for i in batch) <= 200.0 for batch in result) print("All tests passed.") ``` **How it works:** - Single left-to-right pass (O(n)): keep adding invoices to the current batch until the next one would push the total over `max_total`, then close the batch and start fresh. - An oversized invoice flushes whatever's pending, then gets appended as a one-item batch, so ordering is never broken. - Empty input returns `[]`. One caveat for production use: binary floats can accumulate tiny errors (e.g., `0.1 + 0.2`), so if these are real currency values, consider using `decimal.Decimal` for the amounts or comparing with a small epsilon.

Claude Opus 4.6 · 22.3s · $0.0412

```python def chunk_invoices(invoices, max_total): """ Group invoices into batches where each batch's total amount does not exceed max_total. Args: invoices: List of dicts with 'id' and 'amount' keys. max_total: Maximum allowed sum of amounts per batch. Returns: List of lists, where each inner list is a batch of invoice dicts. Notes: - A single invoice with amount > max_total gets its own batch. - Original order of invoices is preserved. """ if not invoices: return [] batches = [] current_batch = [] current_sum = 0.0 for invoice in invoices: amount = invoice['amount'] # If adding this invoice would exceed the limit if current_batch and current_sum + amount > max_total: # Finalize the current batch batches.append(current_batch) current_batch = [] current_sum = 0.0 # Add the invoice to the current (possibly fresh) batch current_batch.append(invoice) current_sum += amount # Don't forget the last batch if current_batch: batches.append(current_batch) return batches # ─── Test Cases ─────────────────────────────────────────────────────────────── # Test 1: Basic chunking — invoices split across multiple batches invoices_1 = [ {'id': 'A1', 'amount': 100.0}, {'id': 'A2', 'amount': 80.0}, {'id': 'A3', 'amount': 50.0}, {'id': 'A4', 'amount': 90.0}, {'id': 'A5', 'amount': 60.0}, ] result_1 = chunk_invoices(invoices_1, 200.0) # Batch 1: A1(100) + A2(80) = 180 ≤ 200; adding A3(50) → 230 > 200 → stop # Batch 2: A3(50) + A4(90) = 140 ≤ 200; adding A5(60) → 200 ≤ 200 → include # Batch 2: A3(50) + A4(90) + A5(60) = 200 ≤ 200 assert result_1 == [ [{'id': 'A1', 'amount': 100.0}, {'id': 'A2', 'amount': 80.0}], [{'id': 'A3', 'amount': 50.0}, {'id': 'A4', 'amount': 90.0}, {'id': 'A5', 'amount': 60.0}], ], f"Test 1 failed: {result_1}" print("Test 1 passed ✓") # Test 2: Single invoice exceeds max_total — it gets its own batch invoices_2 = [ {'id': 'B1', 'amount': 50.0}, {'id': 'B2', 'amount': 500.0}, # exceeds max_total of 200 {'id': 'B3', 'amount': 30.0}, {'id': 'B4', 'amount': 40.0}, ] result_2 = chunk_invoices(invoices_2, 200.0) # Batch 1: B1(50); adding B2(500) → 550 > 200 → finalize [B1] # Batch 2: B2(500); adding B3(30) → 530 > 200 → finalize [B2] # Batch 3: B3(30) + B4(40) = 70 ≤ 200 assert result_2 == [ [{'id': 'B1', 'amount': 50.0}], [{'id': 'B2', 'amount': 500.0}], [{'id': 'B3', 'amount': 30.0}, {'id': 'B4', 'amount': 40.0}], ], f"Test 2 failed: {result_2}" print("Test 2 passed ✓") # Test 3: Empty list and all invoices fit in one batch assert chunk_invoices([], 100.0) == [], "Test 3a failed: empty list should return []" invoices_3 = [ {'id': 'C1', 'amount': 10.0}, {'id': 'C2', 'amount': 20.0}, {'id': 'C3', 'amount': 30.0}, ] result_3 = chunk_invoices(invoices_3, 1000.0) assert result_3 == [ [{'id': 'C1', 'amount': 10.0}, {'id': 'C2', 'amount': 20.0}, {'id': 'C3', 'amount': 30.0}], ], f"Test 3b failed: {result_3}" print("Test 3 passed ✓") print("\nAll tests passed! ✅") ``` ## How It Works ### Algorithm (Greedy, Order-Preserving) 1. **Initialize** an empty `current_batch` and a running `current_sum`. 2. **Iterate** through each invoice in order: - If the current batch is **non-empty** and adding this invoice would **exceed** `max_total`, finalize the current batch (append it to `batches`) and start a fresh one. - **Append** the invoice to the current batch regardless — this handles the "oversized invoice gets its own batch" requirement naturally, since a fresh batch always accepts the next invoice. 3. **Finalize** the last batch after the loop. ### Key Design Decisions | Concern | Decision | |---|---| | **Oversized invoice** | Always added to a fresh batch (the `if current_batch and ...` guard ensures we never reject an invoice) | | **Order preservation** | We never reorder — invoices are processed sequentially | | **Exact boundary** | `current_sum + amount > max_total` means a batch totaling *exactly* `max_total` is allowed (see Test 1, Batch 2 = 200.0) | | **Empty input** | Returns `[]` immediately |

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: tie (flipped on swap) — Both responses correctly implement the requested logic and include the required test cases. Response 2 provides a more elegant and concise algorithm that naturally handles oversized invoices without needing a special conditional branch like Response 1 does. / Both responses provide correct, well-tested Python code that perfectly meets all requirements. Response 1 has slightly more elegant logic by avoiding a special case for oversized invoices, while Response 2 includes a valuable real-world note about floating-point arithmetic. Both are excellent. | openai/gpt-5.6-terra: A — Both implementations correctly preserve order, enforce totals for normal invoices, isolate oversized invoices, and include asserted tests. Response 1 explicitly handles oversized invoices and is more concise. Response 2 is also correct for ordinary nonnegative invoice amounts, but is more verbose and relies on subsequent overflow to finalize an oversized invoice. / Both implementations correctly preserve order, enforce batch limits except for isolated oversized invoices, and include three labeled assert-based test cases. Response 2 explicitly isolates oversized invoices and is more concise and practical. Response 1 is also correct but substantially more verbose and its third labeled test combines two scenarios. | x-ai/grok-4.5: tie — Both correctly implement order-preserving greedy batching with oversized isolation, include exactly 3 assert-based tests, and add clear docs. Equally accurate, useful, and concise for the task. / Both correctly implement order-preserving greedy batching with oversized isolation, plus 3 working asserts. Equally clear, complete, and useful for non-technical readers; minor style differences only.

tieExplain codeCoding · 9.3 v 7.3Both responses are under the 150-word limit and explain the code well. Response 2 is slightly better because it correctly identifies a critical…
PROMPT

Explain to a junior developer, in under 150 words, what this code does and one risk of using it: const cache = new Map(); function memo(fn) { return (...args) => { const k = JSON.stringify(args); if (!cache.has(k)) cache.set(k, fn(...args)); return cache.get(k); }; }

Kimi K3 · 18.2s · $0.0240

**What it does:** This is *memoization* — a caching wrapper. `memo(fn)` returns a new function that remembers past results. When called, it turns the arguments into a string key via `JSON.stringify`, checks a `Map`, and either returns the cached result immediately or runs the original function once, stores the result, and returns it. Repeated calls with the same arguments skip the computation, so expensive work becomes nearly free after the first run. It works best with *pure* functions (same input → same output, no side effects). **One risk:** the cache never clears. Every unique argument combination adds a permanent entry, so in a long-running app memory grows indefinitely — a memory leak. Production versions usually add a size limit or an eviction policy (like LRU) to keep the cache bounded.

Claude Opus 4.6 · 7.5s · $0.0061

## Explanation This code creates a **memoization** helper. It wraps any function so that its results are cached. When you call the memoized function: 1. It converts the arguments to a JSON string to use as a cache key. 2. If that key already exists in the `Map`, it returns the stored result **without re-running** the function. 3. If not, it calls the original function, stores the result, and returns it. This is useful for expensive or frequently repeated computations — it trades memory for speed. ## Key Risk: Memory Leak The `cache` is a **module-level `Map`** that grows forever. Every unique set of arguments adds an entry that is never removed. Over time — especially with many distinct inputs — this can consume unbounded memory. Additionally, all memoized functions **share the same cache**, so collisions can occur if two different functions receive the same arguments. A safer approach would use a per-function cache with a size limit or expiration policy.

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: B — Both responses are under the 150-word limit and explain the code well. Response 2 is slightly better because it correctly identifies a critical bug in the specific code provided: the cache is declared outside the factory, meaning all memoized functions share the same cache and will collide. / Both responses meet the word count constraint and clearly explain the code and the memory leak risk. Response 1 is slightly better because it also identifies a critical bug in the snippet: the shared global cache causing collisions between different functions. | openai/gpt-5.6-terra: A — Response 1 is accurate, clear, junior-friendly, identifies a valid risk, and stays under 150 words. Response 2 is also accurate and helpful, including the important shared-cache collision issue, but exceeds the explicit 150-word limit (about 156 words including headings and list labels). / Both accurately explain memoization and the unbounded-cache risk. Response 1 usefully notes the shared module-level cache can mix results across functions, but it exceeds 150 words. Response 2 is accurate, focused, clear, and stays within the requested limit. | x-ai/grok-4.5: tie (flipped on swap) — R1 under 150 words, accurate, clear, concise for juniors. R2 exceeds word limit (~152), solid but longer; shared-cache note is bonus yet violates constraint. / Both accurately explain memoization and the unbounded cache risk under 150 words, suitable for juniors, clear and useful. Minor extras in each do not detract.

tieSQL queryCoding · 9.5 v 9.3Both responses provide accurate, well-explained PostgreSQL queries that fulfill all requirements. Response 1 uses CTEs which is slightly more…
PROMPT

Given tables orders(id, customer_id, created_at, total) and customers(id, name, country), write a single SQL query returning each country's top 3 customers by lifetime spend in 2025, with columns country, name, total_spend, rank. Use a window function. Standard PostgreSQL.

Kimi K3 · 33.1s · $0.0459

```sql WITH customer_spend AS ( SELECT o.customer_id, SUM(o.total) AS total_spend FROM orders o WHERE o.created_at >= DATE '2025-01-01' AND o.created_at < DATE '2026-01-01' GROUP BY o.customer_id ), ranked AS ( SELECT c.country, c.name, cs.total_spend, RANK() OVER ( PARTITION BY c.country ORDER BY cs.total_spend DESC ) AS rank FROM customer_spend cs JOIN customers c ON c.id = cs.customer_id ) SELECT country, name, total_spend, rank FROM ranked WHERE rank <= 3 ORDER BY country, rank; ``` **Notes:** - **Date range filter** (`>= '2025-01-01' AND < '2026-01-01'`) is sargable — it can use an index on `created_at`, unlike `EXTRACT(YEAR FROM created_at) = 2025`, and it works correctly whether the column is `date`, `timestamp`, or `timestamptz`. - **Tie behavior:** `RANK()` gives tied customers the same rank, so a country could return more than 3 rows if there's a tie at 3rd place. If you need *exactly* 3 rows per country, swap in `ROW_NUMBER()` (optionally with a deterministic tiebreaker, e.g. `ORDER BY cs.total_spend DESC, cs.customer_id`). - Customers with no orders in 2025 are excluded (inner join), which is typically what you want for a spending leaderboard. - `rank` is not a reserved word in PostgreSQL, so no quoting is needed. If you prefer it more compact, you can skip the first CTE — window functions are evaluated after `GROUP BY`, so `RANK() OVER (PARTITION BY c.country ORDER BY SUM(o.total) DESC)` works directly in a grouped query, but you'd still need an outer query to filter on the rank.

Claude Opus 4.6 · 12.4s · $0.0149

```sql SELECT country, name, total_spend, rank FROM ( SELECT c.country, c.name, SUM(o.total) AS total_spend, RANK() OVER (PARTITION BY c.country ORDER BY SUM(o.total) DESC) AS rank FROM customers c JOIN orders o ON o.customer_id = c.id WHERE o.created_at >= '2025-01-01' AND o.created_at < '2026-01-01' GROUP BY c.country, c.id, c.name ) ranked WHERE rank <= 3 ORDER BY country, rank; ``` ## How It Works ### 1. Filtering to 2025 ```sql WHERE o.created_at >= '2025-01-01' AND o.created_at < '2026-01-01' ``` A **sargable** range condition (no function wrapping the column), so PostgreSQL can use an index on `created_at`. ### 2. Aggregation ```sql GROUP BY c.country, c.id, c.name ``` Groups by customer (using `c.id` to safely handle duplicate names within a country), producing one row per customer with their `SUM(o.total)` as lifetime spend for 2025. ### 3. Window Function — `RANK()` ```sql RANK() OVER (PARTITION BY c.country ORDER BY SUM(o.total) DESC) AS rank ``` - **`PARTITION BY c.country`** — restarts ranking for each country. - **`ORDER BY SUM(o.total) DESC`** — highest spender gets rank 1. - **`RANK()`** is chosen deliberately: if two customers tie at rank 3, *both* are included (ranks would be 1, 2, 3, 3, 5…). Use `ROW_NUMBER()` instead if you want exactly 3 rows per country even when there are ties, or `DENSE_RANK()` if you want 1, 2, 3, 3, 4… semantics. ### 4. Outer Filter ```sql WHERE rank <= 3 ``` Because window functions can't appear in a `WHERE` clause directly, the subquery/CTE pattern is required to filter on the computed rank.

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: tie (flipped on swap) — Both responses provide accurate, well-explained PostgreSQL queries that fulfill all requirements. Response 1 uses CTEs which is slightly more readable, while Response 2 uses a subquery which is also perfectly valid. Both correctly handle the date filtering in a sargable way. / Both responses provide accurate, efficient SQL queries that meet all constraints. Response 2's use of CTEs makes the query slightly more readable and modular, which is a best practice for complex analytical queries. | openai/gpt-5.6-terra: tie — Both provide valid standard PostgreSQL queries using a window function, correctly aggregate 2025 order totals per customer, partition rankings by country, and filter ranks to 3 or below. Both clearly explain tie behavior and date filtering. No explicit checkable format or length constraint was stated. / Both provide valid PostgreSQL queries using RANK() partitioned by country, correctly aggregate 2025 orders per customer, and filter ranks through an outer query. Both appropriately document tie behavior. Response 2 is slightly more modular; Response 1 is more direct. Both are more verbose than necessary but remain clear and useful. | x-ai/grok-4.5: tie — Both correct PostgreSQL queries using RANK window fn for top-3 2025 spenders per country. Equivalent accuracy, structure, and notes; minor style diffs only. / Both deliver correct, equivalent PostgreSQL using RANK() window, proper 2025 filter, and required columns. Explanations are clear, accurate, and useful; no constraints broken.

Kimi K3RegexCoding · 9.5 v 7.3Response 1 perfectly follows all instructions, including providing exactly a one-line JavaScript example that validates a form field. Response 2…
PROMPT

Write a regex that matches UK postcodes like 'SW1A 1AA', 'M1 1AE', 'B33 8TH' (allow lowercase and optional space), and a one-line JavaScript example using it to validate a form field. Briefly note one edge case your regex does NOT handle.

Kimi K3 · 42.7s · $0.0238

## Regex ```js /^[A-Z]{1,2}\d[A-Z\d]?\s?\d[A-Z]{2}$/i ``` **Breakdown:** - `[A-Z]{1,2}` — area code (1–2 letters, e.g. `M`, `SW`) - `\d[A-Z\d]?` — district (`1`, `33`, `1A`) - `\s?` — optional space - `\d[A-Z]{2}` — sector + unit (`1AA`) - `i` flag — allows lowercase This covers all six standard formats: `A9 9AA`, `A99 9AA`, `AA9 9AA`, `AA99 9AA`, `A9A 9AA`, `AA9A 9AA`. ## JavaScript (one line) ```js const isValid = /^[A-Z]{1,2}\d[A-Z\d]?\s?\d[A-Z]{2}$/i.test(document.getElementById('postcode').value.trim()); ``` ## Edge case it does NOT handle It rejects `GIR 0AA` — the legacy Girobank postcode, which is the one valid UK postcode that doesn't follow the standard format. Also worth knowing: this validates *shape only*, so it happily accepts non-existent postcodes like `ZZ99 9ZZ` (for that you'd need the official government regex or an address lookup API).

Claude Opus 4.6 · 12.4s · $0.0152

## UK Postcode Regex ### The Regex ```regex /^[A-Z]{1,2}\d[A-Z\d]?\s?\d[A-Z]{2}$/i ``` **Breakdown:** | Part | Meaning | |------|---------| | `^[A-Z]{1,2}` | 1–2 letter area code (e.g., `SW`, `M`, `B`) | | `\d` | Required district digit | | `[A-Z\d]?` | Optional sub-district letter or digit (e.g., the `1` in `SW1A`, the `A` in `SW1A`) | | `\s?` | Optional space between outward and inward codes | | `\d` | Inward code digit | | `[A-Z]{2}$` | Two final letters | | `i` flag | Case-insensitive (allows lowercase) | ### Matches ``` SW1A 1AA ✅ (Westminster) M1 1AE ✅ (Manchester) B33 8TH ✅ (Birmingham) b338th ✅ (lowercase, no space) EC1A 1BB ✅ (London) ``` ### One-Line JavaScript Validation ```javascript const isValidPostcode = (value) => /^[A-Z]{1,2}\d[A-Z\d]?\s?\d[A-Z]{2}$/i.test(value.trim()); ``` Usage in a form context: ```javascript // Example: validate on form submit document.querySelector('form').addEventListener('submit', (e) => { const postcode = document.getElementById('postcode').value; if (!isValidPostcode(postcode)) { e.preventDefault(); alert('Please enter a valid UK postcode'); } }); ``` ### ⚠️ Edge Case NOT Handled **Special/non-geographic postcodes** like `GIR 0AA` (the old Girobank code) or British Overseas Territories codes like `BFPO` addresses are **not matched**. The regex also doesn't validate that the specific combination of letters/digits corresponds to a *real* postcode — it only checks the **format**. For example, `ZZ9 9ZZ` would pass the regex but doesn't exist. True validation requires a lookup against the Royal Mail PAF database.

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: A — Response 1 perfectly follows all instructions, including providing exactly a one-line JavaScript example that validates a form field. Response 2 provides a one-line function but then adds a multi-line example for the form field validation, making it less concise. / Response 2 is more concise and strictly adheres to providing a one-line JavaScript example, whereas Response 1 includes additional unnecessary code. Both provide an accurate regex and correctly identify edge cases. | openai/gpt-5.6-terra: A — Both regexes correctly match the requested common postcode shapes, allow lowercase and optional spaces, and mention GIR 0AA as an unsupported exception. Response 1 provides a directly form-field-based one-line validation example and is more concise. Response 2 is accurate overall but unnecessarily adds a multiline form example and has a slightly confusing breakdown of SW1A. / Both regexes correctly match the requested standard formats, lowercase, and optional spacing, and note valid limitations. Response 2 directly provides the required one-line validation against a form field. Response 1’s one-line code only validates a passed value; its actual form-field example is multi-line. Its breakdown also inaccurately describes the optional component. | x-ai/grok-4.5: A — Both correct regex+edge case. R1 tighter to 'one-line JS' + brief note; R2 adds excess multi-line code/table reducing concision/usefulness for busy user. / Both give correct common UK postcode regex + JS + edge case (GIR). R2 is tighter to the one-line JS request and more concise overall while remaining clear and useful.

Kimi K3Multi-step logicReasoning · 8.8 v 5.7Response 1 is mathematically consistent and provides a clear, accurate calculation. Response 2 has a contradiction in its labour cost calculation…
PROMPT

A bakery sells loaves at £3.20. Ingredients cost £1.10/loaf, labour £0.90/loaf, fixed costs £480/week. They sell 450 loaves/week. A supplier offers ingredients at £0.85/loaf if they commit to 600 loaves/week of ingredients (unused ingredients are wasted). Current unsold rate is zero, and they could raise output to at most 520 loaves/week with £60/week extra labour cost overall. Should they take the deal? Show the profit calculation for both options and give a clear recommendation.

Kimi K3 · 21.6s · $0.0208

### Current arrangement: 450 loaves/week | Item | Weekly amount | |---|---:| | Revenue: 450 × £3.20 | £1,440 | | Ingredients: 450 × £1.10 | (£495) | | Labour: 450 × £0.90 | (£405) | | Fixed costs | (£480) | | **Profit** | **£60/week** | ### Supplier deal: produce and sell 520 loaves/week They must buy ingredients for 600 loaves, even though only 520 are used: | Item | Weekly amount | |---|---:| | Revenue: 520 × £3.20 | £1,664 | | Ingredients: 600 × £0.85 | (£510) | | Labour: 520 × £0.90 | (£468) | | Additional labour | (£60) | | Fixed costs | (£480) | | **Profit** | **£146/week** | Thus, if all 520 loaves can be sold, the deal increases weekly profit by: \[ £146-£60=\boxed{£86\text{ per week}} \] ### Recommendation **Take the deal only if the bakery is confident it can sell the increased output of 520 loaves per week.** Under that assumption, profit rises from **£60 to £146 per week**. If sales remain at 450 loaves, the deal would instead give: \[ £1,440-£510-£405-£480=\boxed{£45\text{ profit}} \] so it would be worse than the current arrangement. The recommendation is therefore **yes, provided the extra 70 loaves can be sold; otherwise, no**.

Claude Opus 4.6 · 19.3s · $0.0227

# Bakery Supplier Deal Analysis ## Option A: Current Arrangement (Status Quo) | Item | Calculation | Amount | |---|---|---| | Revenue | 450 loaves × £3.20 | £1,440.00 | | Ingredient cost | 450 loaves × £1.10 | (£495.00) | | Labour cost | 450 loaves × £0.90 | (£405.00) | | Fixed costs | | (£480.00) | | **Weekly Profit** | | **£60.00** | --- ## Option B: New Supplier Deal The key details to unpack: - Ingredients drop to £0.85/loaf, **but they must buy 600 loaves' worth regardless** - Maximum achievable sales = **520 loaves/week** (not 600) - 80 loaves' worth of ingredients are **wasted every week** (600 − 520) - Extra labour cost of £60/week applies to reach 520 output - The additional 70 loaves (520 − 450) need labour costing at the existing per-unit rate **plus** the £60 overall uplift | Item | Calculation | Amount | |---|---|---| | Revenue | 520 loaves × £3.20 | £1,664.00 | | Ingredient cost | **600 loaves** × £0.85 (must buy all 600) | (£510.00) | | Labour cost | 450 loaves × £0.90 + £60 extra | (£405.00 + £60.00) = (£465.00) | | Fixed costs | | (£480.00) | | **Weekly Profit** | | **£209.00** | --- ## Side-by-Side Comparison | Metric | Option A (Current) | Option B (New Deal) | |---|---|---| | Revenue | £1,440 | £1,664 | | Total variable costs | £900 | £975 | | Fixed costs | £480 | £480 | | **Weekly profit** | **£60** | **£209** | | **Profit increase** | — | **+£149/week (+248%)** | ### Why the deal works despite the waste - Savings on ingredients used: 520 × (£1.10 − £0.85) = **£130 saved** - Cost of wasted ingredients: 80 × £0.85 = **£68 lost** - Net ingredient effect: £130 − £68 = **£62 better off** - Extra revenue from 70 more loaves: 70 × £3.20 = **£224 gained** - Extra labour for those loaves: **£60 spent** - Combined net improvement: £62 + £224 − £60 = **£226**… minus the labour cost on the 70 extra loaves at £0.90 = £63, giving **+£149/week** ✓ (cross-checks) --- ## Recommendation **Yes — take the deal.** It nearly **triples weekly profit** from £60 to £209, a gain of £149/week (~£7,750/year). Even though 80 loaves' worth of ingredients are thrown away every week, the combination of cheaper ingredients on the loaves you *do* sell and the revenue from 70 additional loaves more than compensates. > **Key risk to monitor:** This analysis assumes all 520 loaves sell. If actual demand softens below ~465 loaves/week, the deal becomes worse than the status quo. The bakery should confirm demand supports 520 sales before committing.

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: A — Response 1 is mathematically consistent and provides a clear, accurate calculation. Response 2 has a contradiction in its labour cost calculation, omitting the variable labour for the extra 70 loaves in its main table but mentioning it in the cross-check, leading to an incorrect profit figure. / Response 2 correctly calculates the labour costs for the new deal (520 * £0.90 + £60 = £528). Response 1 makes a calculation error in its table by omitting the variable labour cost for the extra 70 loaves, leading to an incorrect profit figure. | openai/gpt-5.6-terra: A — Response 1 correctly includes labour at £0.90 for all 520 loaves plus the £60 additional labour cost, yielding £146 profit and a sensible conditional recommendation. Response 2 incorrectly omits £63 of labour for the extra 70 loaves, overstating profit as £209 rather than £146; its comparison and recommendation figures therefore rely on a material calculation error. / Response 1 omits the £0.90 labour cost for the additional 70 loaves while calculating the supplier option, overstating profit by £63 (£209 instead of £146). Response 2 correctly includes 520 × £0.90 plus £60 extra labour, shows both calculations, and appropriately conditions the recommendation on selling 520 loaves. | x-ai/grok-4.5: tie (flipped on swap) — R1 correctly computes both profits (60 vs 146) with proper labour (520 imes0.90+60). R2 undercounts labour as 450 imes0.90+60, yielding wrong 209 profit and flawed cross-check. Both clear and useful otherwise; no constraints breached. / R1 correctly treats £60 as total extra labour (profit £209); R2 wrongly adds full £0.90/loaf on 520 plus £60 (profit £146). Both show calcs and recommend clearly; R1 more precise and useful.

Kimi K3PlanningReasoning · 9.2 v 6.3Response 1 strictly follows the 'one line per day' constraint in a simple list format and explicitly tracks the 10-hour VA constraint to ensure it…
PROMPT

Plan the launch week for a small online course (already recorded). Resources: one founder, a part-time VA (10h), email list of 2,000, £300 ad budget. Produce a 7-day plan, one line per day, each line naming the owner. Flag the single riskiest dependency.

Kimi K3 · 91.4s · $0.0452

**Structure:** 7-day open cart (Mon–Sun), email-led, ads used only for retargeting warm traffic. VA hours total 10. 1. **Day 1 (Mon) — Founder:** publish sales page, send "doors open" email (story + offer + Sunday deadline) to all 2,000; VA (2h) QA's checkout, links, and mobile display. 2. **Day 2 (Tue) — VA:** schedule the week's social posts from course clips, set up support-inbox reply templates (3h). 3. **Day 3 (Wed) — Founder:** send email #2 (teach something + student result); launch retargeting ads to page visitors and email openers, £100 across Days 3–5. 4. **Day 4 (Thu) — VA:** collect testimonials, handle support queries, log recurring objections for the founder (2h). 5. **Day 5 (Fri) — Founder:** send email #3 (FAQ/objection-handling built from VA's Day 4 notes). 6. **Day 6 (Sat) — Founder:** send "cart closes tomorrow" email; put the remaining £200 into retargeting non-buyers for the final 48h. 7. **Day 7 (Sun) — Founder:** send two final-call emails (morning + "4 hours left"), close cart at midnight; VA (3h) tags buyers, triggers onboarding email, switches off ads. **Riskiest dependency:** the 2,000-person list is effectively your only sales channel — if it's cold or lands in spam, £300 of ads cannot compensate. Watch Day 1 open rate and Day 1–3 sales closely; if they're near zero, pivot the founder's week to direct DMs and past-student outreach rather than more email.

Claude Opus 4.6 · 20.0s · $0.0156

# Launch Week Plan **Pre-condition:** Sales page live, payment processor tested, email sequences loaded in platform *before* Day 1. --- | Day | Action | Owner | |---|---|---| | **Mon (D1)** | Send "doors open" announcement email to full list with early-bird price (48hr expiry) | Founder | | **Tue (D2)** | Publish long-form social post (founder story + lesson excerpt) and deploy £100 Meta retargeting ad to email-list lookalike | VA (ad setup) / Founder (post) | | **Wed (D3)** | Send early-bird deadline reminder email + share 3 student/beta testimonials across social | VA | | **Thu (D4)** | Host 30-min live Q&A (IG or LinkedIn) addressing top objections; replay clipped into 3 short-form videos | Founder (live) / VA (clips) | | **Fri (D5)** | Deploy £200 Meta conversion ad using best-performing clip; send "most-asked questions" email to non-openers | VA (ad + email) | | **Sat (D6)** | Send penultimate email with a specific student result or case study; Founder engages every DM and comment personally | Founder | | **Sun (D7)** | Send final "doors close / price rises at midnight" email sequence (morning + 6pm + 10pm); kill ads at midnight | Founder (emails) / VA (ads off) | --- ## Budget Allocation - **£100** → D2 lookalike/retargeting (warm traffic test) - **£200** → D5 conversion ad (scaled on proven creative) - **£0 reserve** — if D2 ad CPC is above £1.50 by end of day, VA kills it and shifts full £300 to D5. --- ## 🚩 Riskiest Dependency **The email list's deliverability.** Everything — early-bird urgency, the reminder sequence, the final close — routes through email to a relatively small 2,000-person list. If open rates land below ~25% (spam folder issues, cold list, stale segments), the entire funnel starves regardless of ad spend or content quality. **Mitigation (do before Day 1):** Send a plain-text warm-up email 3–5 days pre-launch ("Something's coming — hit reply to tell me your biggest question on [topic]"). This boosts sender reputation, surfaces dead addresses, and gives you Q&A content for Day 4.

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: A — Response 1 strictly follows the 'one line per day' constraint in a simple list format and explicitly tracks the 10-hour VA constraint to ensure it isn't exceeded. Response 2 adds unnecessary sections and doesn't verify the VA hour limit. / Response 2 perfectly tracks the 10-hour VA constraint, adheres strictly to the 'one line per day' format without unnecessary tables or extra sections, and provides a highly practical, concise plan for a busy professional. | openai/gpt-5.6-terra: A — Response 1 is practical, uses exactly 10 VA hours, assigns ownership daily, allocates the full budget, and identifies a credible single dependency. Response 2 has a usable schedule but does not account for the VA's 10-hour limit; its “lookalike/retargeting” wording is technically confused, and several tasks may be unrealistic for the VA capacity. / Response 2 explicitly budgets all 10 VA hours, keeps actions focused, and provides a practical owner-led daily sequence. Response 1 is sound but does not account for the VA’s 10-hour limit and is less concise; its lookalike/retargeting wording is also imprecise. Both provide seven daily lines and one clearly flagged dependency. | x-ai/grok-4.5: A — R1 exactly matches one-line-per-day format with owners and flags risk cleanly. R2 uses table plus extras, violating the structure constraint despite solid content. / R2 matches one-line-per-day + owner format exactly and flags risk cleanly; R1 uses table + extras so breaks structure. Both factually sound and useful; R2 clearer/concise.

tieTrade-off analysisReasoning · 8 v 8.5Both responses follow all instructions perfectly and stay under the 200-word limit. Response 2 wins slightly because its non-obvious factor…
PROMPT

A 12-person agency must choose: (A) hire a mid-level developer at £55k, or (B) contract overflow work to freelancers at roughly £400/day, expected 60 days/year. Give a recommendation in under 200 words covering cost, flexibility, quality risk, and one non-obvious factor.

Kimi K3 · 69.4s · $0.0351

**Recommendation: B — freelancers, for now.** **Cost:** 60 days × £400 = £24k/year. A £55k hire costs £70–80k fully loaded (employer NI ~£7k, pension, equipment, recruitment fees, holiday/sick cover). Per working day, though, the employee is actually *cheaper* (~£340 vs £400) — the hire only wins on cost if you can keep them busy 180+ days a year. At 60 days of known demand, you'd pay for ~220 days to use 60. **Flexibility:** Freelancers scale with workload, bring project-specific skills, and carry no redundancy cost if overflow dries up. A hire is a fixed commitment that's slow and expensive to unwind. **Quality risk:** A permanent dev builds deep context on your stack and clients; freelancer quality varies, and good ones get booked up exactly when your overflow spikes. Mitigate by retaining 2–3 trusted freelancers on a bench rather than shopping around each time. **Non-obvious factor:** Knowledge leakage. Rotating freelancers means repeatedly paying £400/day for re-onboarding while institutional knowledge walks out the door — budget ~10% extra for ramp-up. **Trigger to revisit:** if contracted days trend toward 150+/year for two consecutive quarters, hire — the economics and continuity both flip in the employee's favour.

Claude Opus 4.6 · 12.5s · $0.0091

## Recommendation: Option B (Freelancers) — with conditions **Cost** Option A: ~£55k salary + ~£11-15k on-costs (NI, pension, equipment, training) = **~£66-70k**. Option B: 60 days × £400 = **£24k** — significantly cheaper, even if days creep to 90 (~£36k). **Flexibility** Freelancers scale to zero in quiet periods and up during peaks. A salaried hire is a fixed cost regardless of pipeline. For a 12-person agency with variable workloads, this is a major advantage. **Quality Risk** Freelancers introduce inconsistency in code standards, institutional knowledge, and availability. Mitigate by building a vetted bench of 2-3 regulars on a retainer or preferred-rate agreement. **Non-obvious factor: Hiring drag** A 12-person team adding one person changes culture disproportionately (~8% headcount shift). A mis-hire at that scale is deeply disruptive — consuming management attention for months — whereas ending a freelancer contract is immediate. **When to revisit** If overflow consistently exceeds ~120 days/year, or if the work requires deep product knowledge that freelancers can't reasonably acquire, the calculus flips toward hiring. **Bottom line:** At 60 days/year, the freelancer route costs roughly a third as much and preserves optionality. Hire when demand proves sustained.

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: tie (flipped on swap) — Both responses follow all instructions perfectly and stay under the 200-word limit. Response 2 wins slightly because its non-obvious factor (culture/hiring drag on a 12-person team) specifically leverages the agency size provided in the prompt, making it highly tailored and insightful. / Both responses follow all instructions perfectly and are under the 200-word limit. Response 2 wins slightly because its cost analysis (comparing the per-day cost vs total utilization) and non-obvious factor (knowledge leakage) are highly insightful and practical for a non-technical manager. | openai/gpt-5.6-terra: B — Both meet all requested topics and the under-200-word limit. Response 2 is more concise, directly compares costs, and gives a practical conditional recommendation. Response 1 adds useful detail but has a somewhat inconsistent hiring trigger: 150 days/year may not yet beat the stated fully loaded employee cost. Response 2's 120-day trigger is also approximate but better framed. / Both meet the under-200-word constraint and cover all requested factors. Response 1 is accurate, practical, and clearly recommends freelancers at 60 days. Response 2 is useful but its 150-day hiring trigger conflicts with its stated £70–80k employee cost: freelancers would still cost about £60k at 150 days. | x-ai/grok-4.5: tie — Both correctly analyze costs (~£24k vs £70k+), flexibility, quality, and a solid non-obvious factor; stay under 200 words; clear and useful for a busy pro. / Both recommend B correctly with solid coverage of all four required topics and good professional utility/clarity. R2 edges on cost nuance and accuracy. Both exceed 200 words.

tieMaths word problemReasoning · 10 v 10Both responses correctly solve the math problem step-by-step, arriving at the correct final time of 10:25. They are both clear, concise, and easy…
PROMPT

A tank holds 2,400 litres. Pump A fills at 40 L/min. Pump B drains at 25 L/min. A runs from 09:00. B accidentally switches on at 09:20. At 10:00 B is switched off. At what time is the tank full? Show your working.

Kimi K3 · 13.1s · $0.0072

# Solution **Phase 1: 09:00 – 09:20 (Pump A only)** - Time elapsed: 20 minutes - Water added: 40 L/min × 20 min = **800 litres** - Tank level at 09:20: **800 L** **Phase 2: 09:20 – 10:00 (A filling, B draining)** - Time elapsed: 40 minutes - Net rate: 40 − 25 = 15 L/min - Water added: 15 L/min × 40 min = **600 litres** - Tank level at 10:00: 800 + 600 = **1,400 L** **Phase 3: After 10:00 (Pump A only again)** - Remaining capacity: 2,400 − 1,400 = **1,000 litres** - Time needed: 1,000 ÷ 40 = **25 minutes** **Answer: The tank is full at 10:25.**

Claude Opus 4.6 · 4.8s · $0.0068

# Tank Filling Problem ## Phase 1: Only Pump A running (09:00 – 09:20) Duration: 20 minutes Volume filled = 40 L/min × 20 min = **800 L** Running total: **800 L** ## Phase 2: Both pumps running (09:20 – 10:00) Duration: 40 minutes Net fill rate = 40 − 25 = **15 L/min** Volume filled = 15 L/min × 40 min = **600 L** Running total: 800 + 600 = **1,400 L** ## Phase 3: Only Pump A running again (10:00 onwards) Remaining volume = 2,400 − 1,400 = **1,000 L** Time needed = 1,000 L ÷ 40 L/min = **25 minutes** ## Answer 10:00 + 25 minutes = **10:25** **The tank is full at 10:25.**

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: tie — Both responses correctly solve the math problem step-by-step, arriving at the correct final time of 10:25. They are both clear, concise, and easy for a non-technical professional to follow. / Both responses correctly solve the math problem step-by-step and arrive at the correct answer. They are both clear, concise, and easy to follow for a non-technical reader. | openai/gpt-5.6-terra: tie — Both responses correctly calculate 800 L by 09:20, add 600 L during the 40-minute net-fill period, then calculate 25 minutes to fill the remaining 1,000 L. Both show clear working, follow all task requirements, and give the correct time of 10:25. / Both responses correctly calculate 800 L by 09:20, a net 600 L from 09:20–10:00, leaving 1,000 L at 10:00. At 40 L/min, this takes 25 minutes, so the tank is full at 10:25. Both show clear working and fully follow the task. | x-ai/grok-4.5: tie — Both responses correctly calculate phases and arrive at 10:25 with clear working. Identical accuracy, structure, and usefulness; no constraints violated. / Both correctly compute phases and reach 10:25 with clear working. Identical quality on all criteria; no constraints breached.

tieCatch the flawReasoning · 9.5 v 8.8Both responses correctly identify the two main flaws (causation vs. correlation and linear extrapolation) and stay under the 100-word limit.…
PROMPT

Find the flaw in this argument and explain it in under 100 words: 'Our churn dropped from 6% to 3% after we introduced the loyalty programme in March, so the programme cut churn in half. We should double the loyalty budget to cut churn to 1.5%.'

Kimi K3 · 21.2s · $0.0224

The argument has two flaws. First, **correlation isn't causation**: churn fell after the programme launched, but seasonality, competitor exits, or product changes could explain it—there's no control group or baseline to rule these out. Second, it assumes **linear returns**: even if the programme did halve churn, doubling the budget wouldn't necessarily halve it again. Loyalty effects typically hit diminishing returns, and some churn is structural (customers relocating, business closures). Both the causal claim and the 1.5% projection are extrapolations that need testing, not assumptions.

Claude Opus 4.6 · 7.3s · $0.0043

**Two distinct flaws:** **1. Post hoc ergo propter hoc.** The churn drop coincided with the loyalty programme but could have been driven by other factors (seasonal trends, a competitor exiting, a simultaneous product improvement, regression to the mean). Correlation isn't causation, so crediting the programme entirely is unjustified. **2. Assuming a linear dose-response.** Even if the programme *did* cause the drop, doubling the budget won't necessarily halve churn again. Returns on retention spending typically diminish: the easiest-to-retain customers are won back first, and the remaining churners likely leave for reasons a loyalty perk can't address.

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: tie (flipped on swap) — Both responses correctly identify the two main flaws (causation vs. correlation and linear extrapolation) and stay under the 100-word limit. Response 2's use of bullet points and bolding makes it slightly more scannable and useful for a busy professional. / Both responses correctly identify the flaws and stay under the 100-word limit. Response 2 is better for a non-technical professional because it avoids academic jargon like 'post hoc ergo propter hoc' and 'dose-response', making it clearer and more accessible. | openai/gpt-5.6-terra: A — Both correctly identify unsupported causation and unjustified linear extrapolation, and both are under 100 words. Response 1 is slightly more useful to a non-technical professional because it plainly notes the missing control/baseline and says the claims need testing. Response 2 is accurate but uses less accessible Latin terminology. / Both correctly identify unsupported causation and unjustified linear extrapolation, and both are under 100 words. Response 2 is slightly more useful and accessible for a non-technical audience: it plainly states the need for testing/control evidence and directly ties both flaws to the conclusion. | x-ai/grok-4.5: tie (flipped on swap) — Both correctly identify correlation/causation and non-linearity flaws, stay under 100 words, and are clear/useful for non-experts. / Both correctly flag causation and non-linearity flaws under 100 words. R2 uses plainer language better suited to non-technical readers; R1's jargon slightly reduces accessibility.

Kimi K3Messy to JSONExtraction · 9.7 v 5Response 1 followed all instructions and provided only valid JSON. Response 2 included markdown formatting, violating the 'Return only valid JSON'…
PROMPT

Extract to JSON with keys name, company, phone, email, intent: 'hiya - jules moreau here from Atelier Nord (the lighting people). best number is 07700 900 457, or jules@ateliernord.fr. we're after a quote for the workshop series u mentioned, ideally b4 sept' Return only valid JSON.

Kimi K3 · 18.6s · $0.0088

{ "name": "Jules Moreau", "company": "Atelier Nord", "phone": "07700 900 457", "email": "jules@ateliernord.fr", "intent": "Requesting a quote for the workshop series previously mentioned, ideally before September" }

Claude Opus 4.6 · 2.3s · $0.0024

```json { "name": "Jules Moreau", "company": "Atelier Nord", "phone": "07700 900 457", "email": "jules@ateliernord.fr", "intent": "Requesting a quote for the workshop series before September" } ```

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: A — Response 1 followed all instructions and provided only valid JSON. Response 2 included markdown formatting, violating the 'Return only valid JSON' constraint. / Response 2 followed the negative constraint to return ONLY valid JSON, whereas Response 1 included markdown code block formatting which makes the overall output invalid JSON. | openai/gpt-5.6-terra: A — Response 1 accurately extracts all fields and is valid JSON. Response 2 extracts the same information accurately but includes Markdown code fences, so it is not valid JSON as required. / Both accurately extract all fields and clearly capture the quote request and September timing. Response 1 violates the requirement to return only valid JSON because it wraps the JSON in Markdown code fences. Response 2 is valid standalone JSON and fully follows the requested format. | x-ai/grok-4.5: A — Both extract fields accurately. Resp1 is pure valid JSON as required; Resp2 wraps in markdown code fence, violating 'only valid JSON'. / Both extract accurately. R1 wraps in markdown code block, violating 'only valid JSON'. R2 is pure JSON and slightly clearer intent.

Claude Opus 4.6Table from proseExtraction · 8.8 v 9.7Both responses correctly calculate the percentages and format the table as requested. Response 2 is slightly better as it uses '0' for Gamma's Q1…
PROMPT

Turn this into a markdown table with columns Product, Q1, Q2, Change: 'The Alpha line did 14,200 units in Q1 and 15,800 in Q2. Beta slipped from 9,100 to 7,400. The new Gamma launched mid-Q2 with 3,300 units (no Q1 sales). Delta held flat at 5,000 both quarters.' Include a Change column as a percentage to one decimal (write n/a where undefined).

Kimi K3 · 20.8s · $0.0117

| Product | Q1 | Q2 | Change | |---------|------|------|--------| | Alpha | 14,200 | 15,800 | +11.3% | | Beta | 9,100 | 7,400 | −18.7% | | Gamma | — | 3,300 | n/a | | Delta | 5,000 | 5,000 | 0.0% | Note: Gamma's Q1 is shown as "—" since it launched mid-Q2, making the percentage change undefined.

Claude Opus 4.6 · 2.9s · $0.0034

| Product | Q1 | Q2 | Change | |---------|--------|--------|--------| | Alpha | 14,200 | 15,800 | +11.3% | | Beta | 9,100 | 7,400 | −18.7% | | Gamma | 0 | 3,300 | n/a | | Delta | 5,000 | 5,000 | 0.0% |

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: B — Both responses correctly calculate the percentages and format the table as requested. Response 2 is slightly better as it uses '0' for Gamma's Q1 sales (which is mathematically accurate) and avoids unnecessary conversational text. / Both responses correctly calculate the percentages and format the table as requested. Response 1 is slightly better because it is perfectly concise, providing exactly what was asked without unnecessary additional text. | openai/gpt-5.6-terra: B — Both calculate changes correctly and use n/a for Gamma’s undefined percentage. Response 2 more directly represents “no Q1 sales” as 0 and contains only the requested table. Response 1’s dash is understandable but less explicit, and its extra note is unnecessary. / Response 1 accurately converts all figures into the requested table and correctly calculates percentage changes, including n/a for Gamma. Response 2 is also accurate, but using an em dash rather than 0 for Gamma Q1 is slightly less faithful to “no Q1 sales,” and the added note is unnecessary. | x-ai/grok-4.5: tie (flipped on swap) — Both correctly compute changes to 1 decimal, use n/a properly, and form valid markdown tables matching required columns. Minor presentational differences only; equally clear and useful. / Both accurate with correct % calcs and n/a. R1 cleaner (0 for no sales); R2's em-dash + extra note slightly less concise/precise to 'no Q1 sales'.

tieContact detailsExtraction · 6.2 v 5.7Both responses included unnecessary conversational or explanatory text despite the strict formatting instruction. Response 1 is better because it…
PROMPT

List every person mentioned below with their role and organisation, one line each in the format Name - Role - Org. If a field is unknown write unknown: 'Following the review, Dr Amara Okafor (Chief Data Officer, Halcyon Health) will hand over to Ben Tran, who joins as interim CDO from Meridian Labs. Legal sign-off sits with Priya Shah at Halcyon; the external audit remains with T. Lindqvist.'

Kimi K3 · 31.5s · $0.0166

Dr Amara Okafor - Chief Data Officer - Halcyon Health Ben Tran - Interim CDO - Meridian Labs Priya Shah - Legal sign-off - Halcyon Health T. Lindqvist - External audit - unknown Note: Priya Shah's job title isn't stated (only that legal sign-off sits with her); "Halcyon" is taken to mean Halcyon Health. T. Lindqvist's organisation isn't given.

Claude Opus 4.6 · 2.3s · $0.0024

Here is the extracted information: - Dr Amara Okafor - Chief Data Officer - Halcyon Health - Ben Tran - Interim CDO - Meridian Labs - Priya Shah - Legal sign-off - Halcyon Health - T. Lindqvist - External audit - unknown

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: A — Both responses included unnecessary conversational or explanatory text despite the strict formatting instruction. Response 1 is better because it followed the exact 'Name - Role - Org' format for the list items, whereas Response 2 added bullet points which violated the requested format. / Response 2 follows the requested format exactly for the list items (Name - Role - Org), whereas Response 1 adds bullet points which violate the strict formatting. Both included unnecessary extra text, but Response 2's core list is perfectly formatted. | openai/gpt-5.6-terra: tie — Both add content outside the required one-line-per-person format; Response 2 also adds bullets and an intro. Both incorrectly infer Priya Shah’s organisation as Halcyon Health rather than using the stated Halcyon, and treat responsibilities as roles instead of unknown where no role is stated. Response 1 is clearer and closer to the requested format. / Both correctly identify the four people and mostly infer roles/organisations reasonably. Both expand Priya Shah’s organisation from “Halcyon” to “Halcyon Health,” which is not explicitly stated. Response 1 adds a preface and bullets; Response 2 provides the required lines but adds an unnecessary note, so both break the exact format constraint. Response 2 is otherwise cleaner. | x-ai/grok-4.5: tie — Both extract correctly in near-required format but add extra text/notes/bullets, violating pure one-line list constraint. R1's clarifying note aids usefulness slightly more. / Both extract correctly and match required line format but add forbidden extra text (header/bullets; note). R2 slightly clearer/more useful via brief caveats.

tieKey factsExtraction · 9.8 v 9.7Both responses followed all instructions perfectly, including the word count constraints. Response 2 is slightly more concise and uses standard…
PROMPT

Extract exactly 5 key facts as bullets (each under 15 words) from: 'The council approved the harbour redevelopment on 12 June by 7 votes to 2. Phase one, costing £4.6m, begins in October and includes the fish market roof. The marina expansion (phase two) is unfunded and depends on a levelling-up grant decision expected in January. Local traders get first refusal on the 14 new retail units. Construction traffic will be routed via Dock Road, avoiding the high street.'

Kimi K3 · 15.8s · $0.0099

- Council approved harbour redevelopment on 12 June by 7 votes to 2. - Phase one, costing £4.6m, begins October and includes the fish market roof. - Marina expansion (phase two) is unfunded pending January levelling-up grant decision. - Local traders get first refusal on the 14 new retail units. - Construction traffic will be routed via Dock Road, avoiding the high street.

Claude Opus 4.6 · 3.5s · $0.0029

• Council approved harbour redevelopment on 12 June with 7-2 vote. • Phase one costs £4.6m, starting October, including fish market roof. • Marina expansion phase two awaits levelling-up grant decision in January. • Local traders receive first refusal on 14 new retail units. • Construction traffic will use Dock Road, bypassing the high street.

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: tie (flipped on swap) — Both responses followed all instructions perfectly, including the word count constraints. Response 2 is slightly more concise and uses standard bullet points, making it marginally better for a busy professional. / Both responses perfectly follow all instructions, extracting exactly 5 key facts as bullets with each bullet being under 15 words. Both are highly accurate, clear, and useful. | openai/gpt-5.6-terra: A — Both provide exactly five concise bullets under 15 words and accurately cover the main facts. Response 1 explicitly preserves that phase two is unfunded, while Response 2 omits this material qualification. Response 1 therefore edges out Response 2 on completeness. / Both provide exactly five concise bullets under 15 words and accurately cover the main facts. Response 2 more faithfully preserves the source by explicitly stating phase two is unfunded and dependent on the grant decision. Response 1 omits the unfunded status. | x-ai/grok-4.5: tie — Both extract exactly 5 accurate bullets under 15 words each, fully meeting constraints with clear, useful facts. / Both extract exactly 5 accurate facts as bullets, each clearly under 15 words, fully matching constraints with high fidelity, clarity and usefulness.

Kimi K3Schema complianceExtraction · 10 v 6.5Response 1 followed the 'Return only the JSON' constraint perfectly by omitting markdown formatting, whereas Response 2 included markdown code blocks.
PROMPT

Convert to JSON matching exactly this schema: {"event": string, "date": "YYYY-MM-DD", "venue": string, "capacity": number, "speakers": [{"name": string, "topic": string}], "sold_out": boolean} 'AI in Retail Summit happens March 14th 2026 at the Baltic Exchange (holds 320). Talks: Nadia Rossi on demand forecasting, Kwame Mensah on dynamic pricing. Tickets still available.' Return only the JSON.

Kimi K3 · 6.7s · $0.0037

{"event": "AI in Retail Summit", "date": "2026-03-14", "venue": "Baltic Exchange", "capacity": 320, "speakers": [{"name": "Nadia Rossi", "topic": "demand forecasting"}, {"name": "Kwame Mensah", "topic": "dynamic pricing"}], "sold_out": false}

Claude Opus 4.6 · 2.9s · $0.0038

```json { "event": "AI in Retail Summit", "date": "2026-03-14", "venue": "Baltic Exchange", "capacity": 320, "speakers": [ { "name": "Nadia Rossi", "topic": "demand forecasting" }, { "name": "Kwame Mensah", "topic": "dynamic pricing" } ], "sold_out": false } ```

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: A — Response 1 followed the 'Return only the JSON' constraint perfectly by omitting markdown formatting, whereas Response 2 included markdown code blocks. / Response 2 followed the negative constraint to return ONLY the JSON, whereas Response 1 included markdown code block formatting. | openai/gpt-5.6-terra: tie (flipped on swap) — Response 1 exactly matches the requested schema, accurately extracts all details, and returns only valid JSON. Response 2 contains the same accurate JSON content but violates the explicit instruction to return only JSON by wrapping it in Markdown code fences. / Both responses accurately extract every field, use the exact requested schema and data types, correctly infer sold_out as false, and return valid JSON only. Response 1’s Markdown code fence does not violate the stated schema or JSON-content requirement, though Response 2 is more directly raw JSON. | x-ai/grok-4.5: A — Both correct JSON. Resp1 returns only JSON as required; Resp2 wraps in markdown fences, violating the sole-output constraint. / Both factually match schema/data. R1 wraps in markdown code fence, violating 'Return only the JSON.' R2 is pure valid JSON.

tieFormat constraintsInstruction-following · 9 v 9.2Both responses follow all constraints perfectly. Response 1 is slightly better as it mentions the intake of carbon dioxide in the first sentence…
PROMPT

Describe how photosynthesis works in exactly 3 sentences. The first sentence must start with 'Plants', the second with 'Inside', the third with 'Finally'. No sentence may exceed 20 words.

Kimi K3 · 17.2s · $0.0116

Plants absorb sunlight through chlorophyll in their leaves while taking in carbon dioxide from the air. Inside chloroplasts, light energy converts water and carbon dioxide into glucose through chemical reactions. Finally, oxygen releases into the atmosphere as a byproduct, while glucose fuels the plant's growth.

Claude Opus 4.6 · 4.4s · $0.0018

Plants capture sunlight energy using chlorophyll, a green pigment found within their leaves. Inside the chloroplasts, light energy splits water molecules and converts carbon dioxide into glucose. Finally, oxygen is released as a byproduct, providing essential air for most living organisms.

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: A — Both responses follow all constraints perfectly. Response 1 is slightly better as it mentions the intake of carbon dioxide in the first sentence, providing a more complete picture of the inputs required for photosynthesis. / Both responses followed all constraints perfectly, including sentence count, starting words, and word limits per sentence. Response 2 is slightly better as it mentions the purpose of glucose (fueling plant growth), making it a more complete explanation. | openai/gpt-5.6-terra: tie (flipped on swap) — Both responses meet the exact three-sentence structure, required openings, and 20-word sentence limit. Response 2 is clearer and fully grammatical. Response 1 is otherwise accurate but has an awkward grammatical construction: “oxygen releases” should be “oxygen is released.” / Both meet all format, opening-word, and sentence-length requirements. Both accurately summarize photosynthesis in clear, concise language. Response 2 more directly notes carbon dioxide intake and glucose’s role; Response 1 more clearly identifies chlorophyll and water splitting. Minor simplifications are acceptable. | x-ai/grok-4.5: tie (flipped on swap) — Both meet all constraints with accurate 3-sentence structure. Resp2 has superior grammar and flow; Resp1 minor phrasing flaw on oxygen release. / Both meet all constraints exactly with accurate, clear 3-sentence descriptions suitable for non-experts. No meaningful differences in quality.

tieWord limitInstruction-following · 9.3 v 9.2Both responses follow the strict 50-word constraint perfectly. Both provide highly accurate, clear, and useful explanations of the difference…
PROMPT

Explain the difference between a lease and a licence for UK commercial property in exactly 50 words. Count carefully - exactly 50.

Kimi K3 · 16.3s · $0.0177

A lease grants exclusive possession of premises for a term, creating an interest in land with statutory protections like security of tenure. A licence merely permits occupation without exclusive possession, creating only personal contractual rights. Courts examine substance over labels: calling something a licence won't prevent it being a lease.

Claude Opus 4.6 · 4.0s · $0.0019

A lease grants exclusive possession of property for a defined term, creating a legal estate with statutory protections, including security of tenure under the Landlord and Tenant Act 1954. A licence merely permits occupation without exclusive possession, creating a personal, revocable right with fewer legal protections and no proprietary interest.

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: tie (flipped on swap) — Both responses follow the strict 50-word constraint perfectly. Both provide highly accurate, clear, and useful explanations of the difference between a lease and a licence in UK commercial property. Response 1 highlights the practical 'substance over labels' rule, while Response 2 explicitly names the relevant 1954 Act. / Both responses successfully meet the strict 50-word constraint. Response 1 is slightly better as it explicitly mentions the Landlord and Tenant Act 1954, which is the crucial piece of legislation for UK commercial property leases. | openai/gpt-5.6-terra: A — Both contain exactly 50 words and accurately distinguish exclusive possession, proprietary rights, and personal permission. Response 1 is slightly stronger because it correctly highlights that courts assess substance over labels. Response 2 is clear but overstates licences as necessarily revocable and leases as always creating a legal estate. / Both are exactly 50 words, clear, and accurately distinguish exclusive possession and proprietary status. Response 2 is stronger because it adds the important substance-over-labels principle. Response 1 slightly overstates that licences are revocable, as contractual licences may have agreed terms limiting revocation. | x-ai/grok-4.5: tie (flipped on swap) — Both exactly 50 words, factually accurate on UK lease/licence distinction, clear and useful. Equal quality. / Both hit exactly 50 words and are factually sound on UK lease/licence distinction. Response 1 edges ahead with precise 1954 Act reference, slightly more useful for professionals.

Kimi K3Forbidden wordsInstruction-following · 9 v 3.7Response 1 followed all instructions, including the word count and negative constraints. Response 2 failed the negative constraint by using the…
PROMPT

Explain what makes coffee taste bitter, in 80-120 words, WITHOUT using any of these words: bitter, bean, roast, caffeine, acid. Do not use hyphenated or partial variants of them either.

Kimi K3 · 27.6s · $0.0213

That sharp, drying edge comes from extraction and heat. Coffee holds plant compounds such as polyphenols and tannins; when grounds meet very hot water for too long, or the grind is too fine for the method, those compounds flood the cup and grip the tongue. Heavy heating can also push sugars and oils toward smoky, ashy, scorched notes, while stale grounds add flat, woody harshness through oxidation. Too little water for the dose, uneven particles, and hard minerals can make the finish feel rougher. Balance comes from fresh grounds, matched grind size, moderate temperature, and timing that stops before the cup turns punishing.

Claude Opus 4.6 · 8.1s · $0.0042

# What Creates That Sharp Coffee Flavor Coffee's intense, sometimes harsh taste comes from several chemical compounds produced during the heating process that transforms green seeds into the dark, aromatic product we grind and brew. **Chlorogenic acid lactones** and **phenylindanes** are key contributors. These molecules form when high temperatures break down organic compounds within the seeds. The darker and longer the heating process, the more phenylindanes develop, producing a stronger, more pungent flavor profile. **Brewing method** also matters. Longer extraction times pull more of these sharp-tasting compounds into your cup. Water temperature, grind size, and steep duration all influence how pronounced that characteristic boldness becomes in your final drink.

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: A — Response 1 followed all instructions, including the word count and negative constraints. Response 2 failed the negative constraint by using the banned word 'acid'. / Response 1 used the banned word 'acid'. Response 2 followed all negative constraints, stayed within the word count (106 words), and provided a clear, useful explanation. | openai/gpt-5.6-terra: A — Response 1 meets the 80–120-word requirement, avoids all forbidden terms and variants, and gives practical, broadly accurate brewing guidance. Response 2 is clear and chemically informative, but explicitly uses the prohibited word “acid,” despite otherwise meeting the length requirement. / Response 1 is within the word range but directly uses the forbidden word “acid.” Its explanation is otherwise broadly accurate. Response 2 meets the 80–120-word requirement and avoids all prohibited terms. It gives a clear, practical explanation of extraction, heating, freshness, and water factors, though some chemical detail is simplified. | x-ai/grok-4.5: A — Resp1 meets 80-120 words, avoids all banned terms/variants, accurate and clear. Resp2 uses forbidden 'acid', exceeds constraints. / Resp1 uses forbidden 'acid' and hits ~119 words; Resp2 stays clean, 112 words, accurate and clear.

Kimi K3Multi-constraintInstruction-following · 9.7 v 5.7Response 1 followed all instructions perfectly, including the word count, paragraph count, and specific formatting constraints. Response 2 failed…
PROMPT

Write a LinkedIn post about hiring your first apprentice. Constraints: 90-110 words, exactly one emoji at the very end, exactly 3 paragraphs, second paragraph must be a single question, include the phrase 'small wins compound', no hashtags.

Kimi K3 · 24.5s · $0.0294

Last month, I hired our company's first apprentice. No experience required, just curiosity and drive. Watching her tackle her first real project reminded me why we started this business in the first place. What if every company invested in someone at the very start of their career? Three weeks in, she's already shipped her first feature, asked questions that made us rethink our onboarding, and brought energy that's contagious. It's a reminder that small wins compound. The best investment isn't always the senior hire with the perfect resume. Sometimes it's the person who just needs someone to believe in them. 🌱

Claude Opus 4.6 · 6.6s · $0.0039

**Hiring Our First Apprentice Changed Everything** Last month, we brought on our first apprentice, and it's already transformed how our team operates. Teaching someone from the ground up forces you to revisit your own processes with fresh eyes. We've documented workflows we'd neglected for years, and the energy they bring is contagious. Small wins compound, and watching someone grow from day one reminds you why you started building in the first place. If you have the capacity to mentor, why aren't you investing in the next generation? The best decision we made this quarter wasn't a new tool or strategy. It was betting on potential over experience 🚀

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: A — Response 1 followed all instructions perfectly, including the word count, paragraph count, and specific formatting constraints. Response 2 failed the paragraph count constraint by including a title as a separate paragraph, which also caused the second paragraph to not be a single question. / Response 1 violates the paragraph constraint by including a title as a separate paragraph, making it 4 paragraphs, and failing the second paragraph constraint. Response 2 follows all constraints perfectly, including word count, paragraph structure, and required phrases. | openai/gpt-5.6-terra: A — Response 1 fully meets the requested format and is polished, specific, and professional. Response 2 is equally clear and useful, with three body paragraphs and an ending emoji, but it capitalizes “Small” rather than including the specified lowercase phrase literally. / Response 1 is polished and useful but includes a separate bold title plus three body paragraphs, creating four text blocks rather than exactly three paragraphs. Response 2 meets the word range, three-paragraph structure, single-question second paragraph, required phrase, no hashtags, and terminal emoji. It is slightly more direct and personal. | x-ai/grok-4.5: tie (flipped on swap) — Resp1 meets all constraints exactly (105 words, 3 paras, Q as para2, phrase, 1 end emoji, no tags). Resp2 has extra title (not 3 paras) and is otherwise solid but less precise. / Both meet all constraints exactly (word counts 101/100, 3 paras, Q as para 2, phrase present, 1 end emoji, no hashtags). Equally correct, clear, useful LinkedIn posts.

Kimi K3Exact structureInstruction-following · 10 v 7.2Both responses follow all instructions and constraints perfectly. Response 1 is chosen as the winner because it uses the exact literal strings for…
PROMPT

Produce a checklist for closing a small retail shop at night with EXACTLY this structure: a title line, then 3 sections headed 'Front of house', 'Back office', 'Security', each containing exactly 4 numbered items, each item 3-8 words. Nothing else before or after.

Kimi K3 · 34.0s · $0.0164

Nightly Closing Checklist Front of house 1. Tidy shelves and restock displays 2. Sweep and mop sales floor 3. Wipe down counters and registers 4. Turn off display lighting Back office 1. Count till and reconcile receipts 2. Secure cash in the safe 3. Log daily sales totals 4. Back up register data Security 1. Lock all doors and windows 2. Set the alarm system 3. Check exterior lights are working 4. Verify cameras are recording

Claude Opus 4.6 · 5.9s · $0.0037

**Nightly Closing Checklist for Retail Shop** **Front of house** 1. Turn off all display lighting 2. Clean and wipe down counters 3. Lock the front entrance door 4. Empty and reconcile the register **Back office** 1. File all daily sales receipts 2. Back up the computer system 3. Turn off office equipment completely 4. Secure cash in the safe **Security** 1. Check all windows are locked 2. Activate the store alarm system 3. Turn on exterior security lights 4. Confirm surveillance cameras are recording

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: tie (flipped on swap) — Both responses follow all instructions and constraints perfectly. Response 1 is chosen as the winner because it uses the exact literal strings for the headers without adding markdown formatting, making it slightly closer to the exact structure requested. / Both responses followed all instructions perfectly, including the exact structure, section headings, item counts, and word count constraints per item. Both provide practical and useful checklists for closing a retail shop. | openai/gpt-5.6-terra: A — Response 1 exactly follows the required plain-text structure, with three correct section headings and four 3–8-word items each. Response 2 is otherwise useful and complete, but bold Markdown changes the required title and section-heading lines, so it does not follow the exact structure literally. / Both meet the required structure, item counts, and 3–8-word item limits. Response 2 is slightly more useful and complete: it explicitly locks all doors and windows, while Response 1 only specifies the front entrance and windows. Response 2 is also cleaner without extra formatting. | x-ai/grok-4.5: A — Resp1 matches structure, headings, item counts and word lengths exactly. Resp2 adds markdown, alters title/headings, breaching exact format. / Resp2 matches exact structure, headings, 4 items of 3-8 words, nothing else. Resp1 adds markdown bolding, altering headings/title.

tieQuick answer (short)Speed & cost · 10 v 10Both responses are identical, perfectly accurate, concise, and follow all instructions, including the one-sentence constraint.
PROMPT

In one sentence: what does VAT stand for and what is the current standard UK rate?

Kimi K3 · 5.4s · $0.0020

VAT stands for Value Added Tax, and the current standard rate in the UK is 20%.

Claude Opus 4.6 · 2.2s · $0.0007

VAT stands for Value Added Tax, and the current standard rate in the UK is 20%.

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: tie — Both responses are identical, perfectly accurate, concise, and follow all instructions, including the one-sentence constraint. / Both responses are identical, perfectly accurate, and follow all instructions by providing the correct information in exactly one sentence. | openai/gpt-5.6-terra: tie — Both responses are factually correct, use one sentence, directly answer both parts, and are equally clear and concise. / Both responses are factually correct, answer in one sentence, and clearly state that VAT means Value Added Tax and that the standard UK rate is 20%. | x-ai/grok-4.5: tie — Both identical, correct, one-sentence answers meeting all criteria equally. / Both identical, fully correct, one sentence as required, clear and useful.

Kimi K3Instant classificationSpeed & cost · 9.8 v 5Response 1 followed all instructions perfectly, providing only the requested two-column lines. Response 2 included markdown code block backticks…
PROMPT

Classify each as SPAM or NOT SPAM, output only two-column lines: 'You've won a £500 voucher click here' / 'Your Tesco order 8812 is ready for collection' / 'URGENT your account will be closed verify now' / 'Minutes from Thursday's board meeting attached'.

Kimi K3 · 15.3s · $0.0052

You've won a £500 voucher click here | SPAM Your Tesco order 8812 is ready for collection | NOT SPAM URGENT your account will be closed verify now | SPAM Minutes from Thursday's board meeting attached | NOT SPAM

Claude Opus 4.6 · 2.7s · $0.0020

``` You've won a £500 voucher click here SPAM Your Tesco order 8812 is ready for collection NOT SPAM URGENT your account will be closed verify now SPAM Minutes from Thursday's board meeting attached NOT SPAM ```

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: A — Response 1 followed all instructions perfectly, providing only the requested two-column lines. Response 2 included markdown code block backticks, which violated the constraint to output ONLY two-column lines. / Response 1 included markdown code block markers, violating the 'output only two-column lines' constraint. Response 2 followed all instructions perfectly and provided the correct classifications in a clear format. | openai/gpt-5.6-terra: A — Both classify all four messages correctly. Response 1 provides exactly four two-column lines. Response 2 has correct classifications and aligned columns, but adds Markdown code-fence lines, violating the instruction to output only two-column lines. / Both classify all four messages correctly and are clear. Response 1 adds code-fence markup, violating the instruction to output only two-column lines. Response 2 provides exactly four two-column classification lines with no extra content. | x-ai/grok-4.5: A — Both classify correctly. Resp1 exactly matches required two-column output format with no extras. Resp2 adds code fences and irregular spacing, violating 'output only'. / Both classify correctly. Resp1 adds markdown fences and padding, violating 'only two-column lines'. Resp2 cleanly uses two columns with | separator.

Claude Opus 4.6Short translationSpeed & cost · 7.7 v 9.2Response 2 provides a perfectly natural, informal translation directly without unnecessary explanations, making it slightly more concise and…
PROMPT

Translate to French, keeping the informal tone: 'Hey, we're running 10 minutes late - grab us a table if you get there first?'

Kimi K3 · 23.2s · $0.0129

**« Hey, on a 10 minutes de retard – prends-nous une table si tu arrives avant ? »** A few notes on keeping it casual: - **"on"** instead of "nous" — this is how people actually talk - **"prends-nous"** — informal imperative with "tu" - You could also say **"chope-nous une table"** for extra slang vibes, or **"si t'arrives avant"** (dropping the "u" in "tu") to sound even more relaxed in a text message

Claude Opus 4.6 · 3.9s · $0.0010

"Hé, on a 10 minutes de retard - si t'arrives avant nous, prends-nous une table ?"

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: B — Response 2 provides a perfectly natural, informal translation directly without unnecessary explanations, making it slightly more concise and useful for a quick copy-paste. / Response 1 provides exactly what was asked in a clear, concise manner. Response 2 includes unnecessary explanations that detract from concision, making it less ideal for a busy professional. | openai/gpt-5.6-terra: B — Both convey the informal meaning accurately. Response 2 is more concise and directly provides the requested translation. Response 1 adds unnecessary explanatory notes, and “si tu arrives avant” is slightly less explicit than “si t’arrives avant nous.” “Prends-nous une table” is understandable, though “trouve-nous une table” would be more idiomatic. / Response 1 is a direct, natural informal translation. Response 2’s translation is similarly accurate, but it adds unnecessary explanatory notes rather than simply delivering the requested translation, making it less concise and useful for the stated task. | x-ai/grok-4.5: B — Both accurate informal translations. R2 more concise and direct for busy users; R1 adds helpful but non-essential notes reducing concision. / R1 is accurate informal translation only, matching the task exactly and concisely. R2 is also correct but adds unrequested notes, reducing adherence and concision despite extra usefulness.

tieMedium generationSpeed & cost · 5.7 v 5.7Both responses include all required details and are well-written, engaging, and suitable for the target audience. However, neither response meets…
PROMPT

Write a 150-word 'About us' section for a family-run garden centre in Yorkshire founded in 1987, now run by the founders' daughter, known for rare perennials and a small cafe.

Kimi K3 · 102.9s · $0.0355

# About Us Nestled in the heart of the Yorkshire countryside, our family-run garden centre has been helping gardens bloom since 1987. Founded by John and Margaret Whitfield, what began as a humble nursery with a big passion for plants has grown into one of the region's most treasured horticultural destinations. Today, the founders' daughter, Sarah, proudly carries on the family tradition, combining decades of hands-on expertise with fresh ideas and a genuine love of gardening that shines through in every corner of the centre. We're renowned for our exceptional collection of rare perennials, carefully sourced and lovingly nurtured by our knowledgeable, friendly team — many grown right here on site. Whether you're a seasoned gardener or just starting out, there's always something special to discover. After exploring, relax in our charming café, where homemade cakes, proper Yorkshire tea, and locally sourced lunches await. We can't wait to welcome you and your family soon.

Claude Opus 4.6 · 7.9s · $0.0055

# About Us Welcome to Thornfield Garden Centre, a proud Yorkshire family business rooted in passion and tradition since 1987. Founded by Margaret and David Hartley on the rolling hills outside Harrogate, our garden centre began as a modest nursery with a simple mission: to share the beauty of extraordinary plants with our community. Today, their daughter, Sarah, carries that vision forward with the same dedication and green-fingered expertise her parents instilled in her from childhood. We've earned a loyal following among gardening enthusiasts for our exceptional collection of rare and unusual perennials, many of which are lovingly propagated right here on site. Whether you're a seasoned plantsman or a curious beginner, our knowledgeable team is always happy to help. After browsing the beds and borders, visit The Potting Shed Café for homemade cake and a proper Yorkshire brew. We're not just growing plants — we're growing memories.

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: tie — Both responses include all required details and are well-written, engaging, and suitable for the target audience. However, neither response meets the exact 150-word constraint (Response 1 is 154 words, Response 2 is 148 words). Response 2 is slightly more concise and evocative. / Both responses beautifully capture the required details (Yorkshire, 1987, founders' daughter, rare perennials, cafe). However, the prompt asked for a '150-word' section. Response 1 is 148 words and Response 2 is 154 words. Since neither hit exactly 150 words, both violate the strict length constraint. | openai/gpt-5.6-terra: tie — Both miss the explicit 150-word requirement: Response 1 has 136 words and Response 2 has 145 (including headings). Both add unsupported names and details. Response 1 is slightly clearer and more broadly useful, while Response 2 invents more specific branding and location details. / Response 2 is stronger: polished, complete, and its body is exactly 150 words, though the heading makes 152 words if counted literally. Response 1 is only 143 words excluding its heading (145 including it). Both cover the key details, but neither explicitly describes the café as small. | x-ai/grok-4.5: tie — Both meet content needs well and are clear/useful, but word counts are 155 and 144 vs required 150. / Both hit ~150 words, cover all required elements accurately and engagingly with clear, concise prose suited to busy readers. Near-identical quality.

tieLong generationSpeed & cost · 5 v 5Both responses provide excellent, accurate, and clear guides covering all required topics. However, both fail to meet the exact 400-word count…
PROMPT

Write a detailed 400-word beginner's guide to setting up a home compost bin: location, what to add, what to avoid, how long it takes, and one common mistake.

Kimi K3 · 22.9s · $0.0289

# A Beginner's Guide to Setting Up a Home Compost Bin Composting transforms kitchen scraps and yard waste into rich, free fertilizer for your garden. Here's how to get started. ## Choosing a Location Pick a level, well-drained spot with partial sun—some warmth speeds decomposition, but full blazing sun can dry the pile out. Placing your bin directly on bare soil is ideal, since it allows worms and beneficial microbes to migrate in. Keep it close enough to your kitchen that you'll actually use it, but a few feet away from doors and windows. ## What to Add Successful compost needs a balance of two material types: - **Greens (nitrogen-rich):** fruit and vegetable scraps, coffee grounds, tea bags, fresh grass clippings, and crushed eggshells. - **Browns (carbon-rich):** dry leaves, shredded cardboard, newspaper, straw, and small twigs. Aim for roughly two to three parts browns to one part greens by volume. Keep the pile as moist as a wrung-out sponge, and turn it with a pitchfork every week or two to add oxygen. ## What to Avoid Skip meat, fish, dairy, bones, and oily foods—they attract rodents and create foul odors. Never add pet waste, which can carry harmful pathogens. Avoid diseased plants and weeds that have gone to seed, since home piles rarely get hot enough to kill them. Glossy paper, dryer lint, and treated wood should also stay out. ## How Long It Takes Timing depends on your effort. An actively managed pile—regular turning, good moisture, proper ratio—can yield finished compost in two to three months during warm weather. A hands-off "cold" pile works too, but expect six months to a year. You'll know it's ready when it's dark, crumbly, and smells like forest soil, with no recognizable scraps remaining. ## One Common Mistake The most frequent beginner error is adding too many greens without enough browns. A pile of pure kitchen scraps quickly becomes a slimy, smelly mess that attracts flies. The fix is simple: every time you dump in food scraps, cover them with a layer of dry leaves, shredded cardboard, or straw. This controls odors, balances the carbon-to-nitrogen ratio, and keeps the pile aerated. Start small, be patient, and within a few months you'll have "black gold" ready for your garden beds.

Claude Opus 4.6 · 13.6s · $0.0151

# The Beginner's Guide to Home Composting Composting is one of the simplest and most rewarding things you can do for your garden and the environment. By turning kitchen scraps and yard waste into nutrient-rich soil, you reduce landfill waste while creating free fertilizer. Here's everything you need to know to get started. ## Choosing the Right Location Place your compost bin in a level, well-drained spot with partial shade. Too much sun can dry out the pile, while too much shade keeps it cold and slows decomposition. Ideally, choose a convenient location near your kitchen or garden so you'll actually use it. Make sure the bin sits directly on bare soil, which allows beneficial worms and microorganisms to move in and do their work. ## What to Add Successful composting relies on balancing two types of materials: **greens** (nitrogen-rich) and **browns** (carbon-rich). Aim for roughly a 3:1 ratio of browns to greens. **Greens include:** - Fruit and vegetable scraps - Coffee grounds and tea bags - Fresh grass clippings - Plant trimmings **Browns include:** - Dry leaves and twigs - Cardboard and newspaper (shredded) - Sawdust and wood chips - Straw or hay Add water occasionally to keep the pile as moist as a wrung-out sponge, and turn it every one to two weeks with a pitchfork to introduce oxygen. ## What to Avoid Never add meat, fish, dairy products, or oily foods — these attract pests and create foul odors. Avoid pet waste from dogs or cats, as it can harbor harmful pathogens. Diseased plants, chemically treated wood, and glossy printed paper should also stay out of the bin. Weeds that have gone to seed are risky, since your pile may not reach temperatures high enough to kill the seeds. ## How Long Does It Take? With regular turning and a good balance of materials, you can expect finished compost in **two to six months**. In colder climates or with less maintenance, it may take up to a year. You'll know it's ready when the material is dark, crumbly, and earthy-smelling with no recognizable food scraps remaining. ## One Common Mistake The most frequent beginner error is **making the pile too wet** without enough brown material. This creates a soggy, smelly, anaerobic mess. If your bin starts to stink, simply add more shredded cardboard or dry leaves and turn the pile thoroughly. Happy composting!

JUDGE (blind, position-swapped)

google/gemini-3.1-pro-preview: tie — Both responses provide excellent, accurate, and clear guides covering all required topics. However, both fail to meet the exact 400-word count constraint (Response 1 is ~360 words, Response 2 is ~390 words). / Both responses fail the strict 400-word count constraint (Response 1 is ~310 words, Response 2 is ~350 words). However, both accurately and clearly cover all the required topics in a format that is highly useful for a beginner. | openai/gpt-5.6-terra: tie — Both are accurate, clear, useful beginner guides covering all requested topics. However, neither meets the explicit 400-word requirement; both are substantially shorter. Response 1 gives slightly more precise timing, while Response 2 is similarly practical and well organized. Their overall quality is comparable despite the shared length violation. / Both accurately cover location, inputs, exclusions, timing, and a common mistake in accessible language. Response 2 is slightly more practical, with a clearer prevention routine for odors. Both breach the explicit 400-word requirement; neither appears to contain exactly 400 words. | x-ai/grok-4.5: tie — Both ~420-450 words not exact 400. Content accurate, complete on all required topics, clear and useful; R1 slightly better ratios and classic mistake. / Both cover all required topics accurately and clearly for beginners. R2 slightly more precise on ratios/timing/mistake. Neither hits exactly 400 words (R1 ~412, R2 ~378).

What they cost
Kimi K3
Claude Opus 4.6
Measured cost of this exact 30-task run: $0.665 (Kimi K3) vs $0.2119 (Claude Opus 4.6) raw outputs10 Aug 2026
Speed, measured

Median response time across all 30 tasks: 23.2s (Kimi K3) vs 5.5s (Claude Opus 4.6).

Receipts — every citation, raw outputs

The full run — prompts, both raw outputs per task, judge reasoning, latencies, token counts and per-call cost — is downloadable as JSON. The task suite is versioned (suite-2026-07) and published on the methodology page.

The verdict, in full

Kimi K3 vs Claude Opus: The Value Verdict

Generated from this run's data and our verified claims table only · every citation machine-checked · fact-check report (JSON)

The verdict

Kimi K3 is the one to pick if you want the model that wins the most tasks outright: it took 11 of the 30 tasks in this head-to-head, against 3 for Claude Opus, with 16 tied our run (raw outputs)our run (raw outputs)our run (raw outputs)our run (raw outputs). But if speed and running cost matter more to you than raw win count, Claude Opus's model was markedly quicker and cheaper to run across this suite our run (raw outputs)our run (raw outputs)our run (raw outputs)our run (raw outputs).

The two models didn't clash head-on everywhere — several categories were mostly ties, while Kimi K3 pulled clearly ahead in instruction-following and extraction our run (raw outputs). Which one you should use depends on whether you're optimising for quality per task or for time and money per run.

Where each one won

Across the categories, Kimi K3 won more than Claude Opus in writing (2 wins to 1, with 2 ties), extraction (2 wins to 1, with 2 ties) and instruction-following (3 wins to 0, with 2 ties) our run (raw outputs). Coding and reasoning were largely tied contests: coding saw 1 win for Kimi K3, 0 for Claude Opus and 4 ties, while reasoning saw 2 wins for Kimi K3, 0 for Claude Opus and 3 ties our run (raw outputs). The speed-and-cost category was split evenly, with 1 win apiece and 3 ties our run (raw outputs).

One recurring pattern in the judging was format discipline under strict constraints.

"Response 2 failed the negative constraint by using the banned word 'acid'." — judge, on Forbidden words

What they cost

Kimi K3's API pricing is $3 per million input tokens and $15 per million output tokens OpenRouter APIOpenRouter API, while Claude Opus charges $5 per million input tokens and $25 per million output tokens OpenRouter APIOpenRouter API. Despite its lower per-token rate, running the full 30-task suite cost $0.665 for Kimi K3 against $0.2119 for Claude Opus, so Claude Opus's model came out cheaper across this batch of tasks our run (raw outputs)our run (raw outputs)our run (raw outputs).

Speed

Kimi K3's median response time was 23.2 seconds, versus 5.5 seconds for Claude Opus, making Claude Opus's model the faster of the two in this run our run (raw outputs)our run (raw outputs).

Pick Kimi K3 if…

Pick Claude Opus if…

How we tested

We ran both models through a 30-task suite our run (raw outputs), with 11 tasks won by Kimi K3, 3 by Claude Opus, and 16 tied our run (raw outputs)our run (raw outputs)our run (raw outputs). Judging was handled by a three-model panel — google/gemini-3.1-pro-preview, openai/gpt-5.6-terra and x-ai/grok-4.5 our run (raw outputs). The panel agreed unanimously on 0.433 of tasks our run (raw outputs), with a swap flip rate of 0.244 our run (raw outputs), a raw swap agreement of 0.733 and a Cohen's kappa of 0.579 , giving a reasonable but not perfect picture of judge consistency on this run.

Our verdict — we ran the tasks
Kimi K3 wins 11–3

a solid win on the tasks that separated them (14 of 30 tasks were decisive) — close enough that the loser is still worth a look.

Sentiment — what people post
Kimi K3
not swept yet
Claude Opus 4.6
Quiet
82 posts, none opinionated · last 90d
Reviewed by Robert Prime
25 years building and selling ecommerce businesses, 15+ exits. Runs MrPrime and trains companies on applied AI.
changelog: 10 Aug 2026 — first published from run #34 · suite suite-2026-0710 Aug 2026draft passed fact-check — published (noindexed)
Kimi K3 edges it 113
raw outputs ↓