Battles / Best-value API

Kimi K3 wins 6–3

a solid win on the tasks that separated them (9 of 30 tasks were decisive) — close enough that the loser is still worth a look.

Kimi K3 vs GPT-5.6 Terra · Best-value APImodels in this battle: Kimi K3 vs GPT-5.6 Terra

tested 12 Aug 202630 tasks3 judges from different labs
Show the full workings
last verified 12 Aug 2026suite suite-2026-07judges: anthropic/claude-sonnet-5, google/gemini-3.1-pro-preview, x-ai/grok-4.5 (none of them a contestant)judge protocol judge-2026-08 — constraint cap enforced in code, applied to 32 judge scores in this runnot statistically decisive — Wilcoxon signed-rank on score margins p=0.9223 (n=22); sign test on win counts p=0.5078 — held to our confidence gatejudge swap-consistency (Cohen’s κ) = 0.432moderate agreement κ measures how often a judge reached the same verdict when the two answers were shown in the opposite order, corrected for agreement you would get by chance. A battle below 0.21 does not publish a winner at all; above it, this number is how much of the scoreline survived the swap test rather than a claim that the judges were reliable.

The evidence

Suite-by-suite
Writing
311
Coding
131
Reasoning
140
Extraction
140
Instruction-following
041
Speed & cost
050

blue = Kimi K3 wins · grey = ties · white = GPT-5.6 Terra wins (5 tasks per suite)

Round-by-round — all 30 tasks
Kimi K3Cold emailWriting · 9 v 8.2Both meet constraints, friendly/professional, under 120 words, clear CTA, no buzzwords. Response 1 has more specific brewery detail (duty, taproom…
PROMPT

Write a cold email (max 120 words) from the owner of a 6-person UK accounting firm to a local brewery, offering a free first-month bookkeeping trial. Friendly but professional, one clear call to action, no buzzwords.

Kimi K3 · 62.7s · $0.0387

**Subject:** Free first month of bookkeeping for [Brewery name] Hi [Name], I'm Sarah, owner of [Firm name], a six-person accountancy practice in [Town]. We already look after a few local food and drink businesses, so we understand the brewery side of things — alcohol duty, VAT, taproom payroll, seasonal cash flow. I'd like to offer you a free first month of bookkeeping. We'll bring your records up to date, reconcile everything and give you a clear picture of where you stand. If you're happy, we carry on. If not, you keep the work and owe us nothing. Fancy a quick chat? Just reply to this email and we'll find a time. Best wishes, Sarah [Surname] [Firm name] | [Phone number] *(Word count: 107, excluding subject line)*

GPT-5.6 Terra · 3.0s · $0.0010

**Subject:** A free month of bookkeeping for your brewery Hi [Name], I’m [Your Name], owner of [Firm Name], a six-person accounting firm based in [Town]. We work with local businesses that want clear, reliable bookkeeping without it taking time away from running the business. Breweries in particular have plenty to keep track of, from stock and supplier invoices to payroll and VAT. To show you how we work, I’d like to offer your first month of bookkeeping free, with no obligation to continue afterwards. If you’re open to it, would you be available for a 15-minute call next week to see whether we could help? Best regards, [Your Name] [Firm Name] [Phone Number]

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: tie (flipped on swap) — Both meet constraints, friendly/professional, under 120 words, clear CTA, no buzzwords. Response 1 has more specific brewery detail (duty, taproom payroll) adding credibility and personal touch, giving it a slight edge in usefulness and specificity. / Both meet word limit, are friendly, professional, with clear CTAs and no buzzwords. Response 2 is slightly more specific and personable, but both are equally effective overall. | google/gemini-3.1-pro-preview: A — Both responses follow all instructions and constraints perfectly. Response 1 is slightly better as it includes specific, relevant details for a UK brewery (alcohol duty, taproom payroll) and has a very natural, friendly tone. / Both responses follow all instructions and constraints perfectly. Response 2 is slightly better because it includes specific brewery-related accounting terms (alcohol duty, taproom payroll), making it more tailored and effective as a cold email. | x-ai/grok-4.5: A — Both under 120 words, meet all constraints with clear CTA and no buzzwords. R1 more tailored to brewery/UK specifics and useful; R2 solid but generic. / Both meet max 120 words, one CTA, no buzzwords, friendly-pro tone. R2 more tailored to brewery ops and slightly clearer value, higher usefulness.

Kimi K3Product descriptionWriting · 9.2 v 5Response 1 stays within the 60-80 word range (71 words) and avoids cliches with vivid gym-specific details. Response 2 is only 59 words, violating…
PROMPT

Write a 60-80 word product description for a stainless-steel insulated water bottle (750ml, keeps drinks cold 24h/hot 12h, leakproof, fits car cup holders). Target audience: gym-goers. Avoid cliches like 'stay hydrated in style'.

Kimi K3 · 40.5s · $0.0230

Built for the weight room, not the shelf. This 750ml stainless-steel bottle keeps water ice-cold for 24 hours—through your warm-up, your last set, and the drive home. Want hot coffee before a 6 a.m. session? It holds heat for 12 hours. The leakproof cap survives being thrown sideways into a packed gym bag, and the base fits standard car cup holders. Double-wall insulation means zero condensation rings on the bench beside you.

GPT-5.6 Terra · 2.4s · $0.0005

Power through workouts with this 750ml stainless-steel insulated water bottle. Double-wall insulation keeps cold drinks chilled for up to 24 hours and coffee or tea hot for up to 12. Its leakproof lid handles gym bags and locker-room commutes without spills, while the slim shape fits standard car cup holders. Durable, reusable, and easy to carry between training sessions.

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: A — Response 1 stays within the 60-80 word range (71 words) and avoids cliches with vivid gym-specific details. Response 2 is only 59 words, violating the word count constraint, and is more generic/less engaging despite being accurate and clear. / Response 1 is 59 words, missing the 60-80 word requirement. Response 2 is 72 words, fits range, avoids clichés, vivid and gym-specific. Both accurate and well-written, but Response 2 better follows constraints and is more engaging and specific. | google/gemini-3.1-pro-preview: A — Response 1 perfectly meets the word count constraint (72 words) and delivers a highly engaging, gym-focused description without cliches. Response 2 falls slightly short of the minimum word count constraint (59 words) and is a bit more generic. / Response 1 is 59 words, violating the 60-80 word constraint. Response 2 is 72 words, meeting the constraint. Response 2 also has a much better tone tailored specifically to the gym-goer target audience. | x-ai/grok-4.5: A — Resp1 meets 60-80 words, hits all specs vividly for gym audience without cliches. Resp2 under word count (59), solid but less engaging. / R1 under word count at 59; both accurate, gym-targeted, cliche-free. R2 more vivid and useful.

Kimi K3Summarise messy notesWriting · 9.3 v 8.5Both accurate and complete, exactly 5 bullets. Response 1 is more concise, matching the 20-second-read requirement better; Response 2 is slightly…
PROMPT

Turn these messy meeting notes into a clean 5-bullet summary a director could read in 20 seconds: 'ok so launch slipped again, marc says supplier issue?? maybe 3 wks. sarah wants to push the email campaign anyway but tom thinks wait. budget - we're 8k over, mostly the packaging redo. Q: do we tell retail partners now or after new date confirmed. also NEED to hire the warehouse temp before august rush. next mtg tues.'

Kimi K3 · 9.6s · $0.0093

- **Launch delayed ~3 weeks** — supplier issue, per Marc. - **Email campaign: decision needed** — Sarah wants to proceed as planned; Tom says wait for the new date. - **Budget $8K over**, driven mainly by the packaging redo. - **Retail partner comms: open question** — notify now or wait until the new date is confirmed? - **Action: hire warehouse temp before the August rush.** Next meeting Tuesday.

GPT-5.6 Terra · 5.1s · $0.0008

- **Launch delayed again** due to a supplier issue; current estimate is approximately **three weeks**. - **Marketing decision pending:** Sarah recommends proceeding with the email campaign, while Tom advises waiting until the launch date is confirmed. - **Budget is $8K over plan**, driven primarily by the packaging redesign. - **Retail partner communications** need a decision: notify them now or wait for a confirmed revised launch date. - **Operational priority:** hire a warehouse temporary worker before the August volume increase; next meeting is Tuesday.

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: A — Both accurate and complete, exactly 5 bullets. Response 1 is more concise, matching the 20-second-read requirement better; Response 2 is slightly wordier, reducing quick scannability. / Both accurate, follow 5-bullet format, and are clear. Response 2 is more concise and scannable in 20 seconds, retaining key details like 'per Marc' while being tighter overall. | google/gemini-3.1-pro-preview: A — Both responses accurately summarize the notes into 5 bullets. Response 1 is slightly punchier and more concise, making it better suited for a quick 20-second read by a busy director. / Both responses accurately summarize the notes into exactly 5 bullets. Response 2 is slightly more concise and punchy, making it better suited for a busy director to read in under 20 seconds. | x-ai/grok-4.5: tie (flipped on swap) — Both accurately capture all key points in exactly 5 clean bullets. Response 1 is tighter and more scannable for a 20-second director read; Response 2 is slightly wordier. / Both deliver accurate 5-bullet summaries covering all notes cleanly and usefully for a director. Equal clarity and fidelity; minor style differences only.

GPT-5.6 TerraTone rewriteWriting · 7 v 9.2Both meet word limit and facts, but R2 is cleaner, warmer, and concise without extra meta-commentary. R1 includes unnecessary offer to adjust…
PROMPT

Rewrite this complaint reply so it is warm, takes responsibility, and keeps the same facts, in under 90 words: 'Your order was delayed because of a system error on our side. The refund has been processed and will arrive in 3-5 days. We cannot offer further compensation. Let us know if you have questions.'

Kimi K3 · 13.0s · $0.0102

Here's a warmer version (66 words): --- Thank you so much for your patience, and I'm truly sorry for the delay — this was caused by a system error on our end, and we take full responsibility. Your refund has been processed and should arrive within 3–5 days. While we're unable to offer further compensation, please know how much we value you. If any questions come up, we're always here and happy to help. --- Want me to adjust the tone further — more formal, or more casual?

GPT-5.6 Terra · 1.9s · $0.0005

We’re very sorry for the delay with your order. A system error on our side caused the issue, and we understand how frustrating this must be. Your refund has now been processed and should arrive within 3–5 days. While we’re unable to offer additional compensation, we appreciate your patience and are here to help with any questions.

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: B — Both meet word limit and facts, but R2 is cleaner, warmer, and concise without extra meta-commentary. R1 includes unnecessary offer to adjust tone, reducing conciseness and professionalism. / Both meet word limit and facts, but Response 2 adds meta-commentary and a follow-up question, violating conciseness/format expectations. Response 1 is clean, warm, concise, and directly usable. | google/gemini-3.1-pro-preview: B — Both responses follow all instructions and constraints, including the word count. Response 2 is better because it provides only the requested text without unnecessary conversational filler, making it more useful and concise. / Response 1 directly provides the rewritten text without unnecessary conversational filler. It perfectly meets all constraints, maintains the facts, and adopts a warm, responsible tone within the word limit. Response 2 includes unwanted conversational text. | x-ai/grok-4.5: B — Both keep facts, are warm, take responsibility, and stay under 90 words. R2 is cleaner and more directly usable. R1 adds unnecessary meta-commentary and offer to revise, reducing concision and professionalism. / Both keep facts, stay warm, own the error, and under 90 words. R1 is pure and concise. R2 adds meta text and offer, reducing clarity for a busy pro.

tieHeadlinesWriting · 8.7 v 8.7Both meet constraints and word limits. Response 1 clearly labels categories matching the mix requirement, aiding clarity. Response 2 lacks…
PROMPT

Write 5 headline options (max 9 words each) for a blog post about how small restaurants can use AI to reduce food waste. Mix: 2 practical, 2 curiosity-driven, 1 with a number.

Kimi K3 · 39.4s · $0.0234

Here are 5 headline options: **Practical:** 1. How Small Restaurants Can Cut Food Waste with AI 2. Practical AI Tools to Reduce Restaurant Food Waste **Curiosity-driven:** 3. Is Your Restaurant Throwing Away Profits? AI Can Help 4. What If Your Trash Could Tell You What's Next? **Number:** 5. 5 Ways AI Helps Small Restaurants Reduce Food Waste All are 9 words or fewer. Want me to generate variations on any of these, or adjust the tone (more playful, more formal, SEO-focused)?

GPT-5.6 Terra · 2.0s · $0.0004

1. Use AI to Cut Restaurant Food Waste 2. Smarter Inventory, Less Waste: AI for Restaurants 3. The AI Secret to Reducing Kitchen Waste 4. Can AI Predict What Your Customers Won’t Order? 5. 5 Ways AI Helps Restaurants Waste Less Food

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: tie (flipped on swap) — Both meet constraints and word limits. Response 1 clearly labels categories matching the mix requirement, aiding clarity. Response 2 lacks category labels, making the required mix less explicit, though content is fine. Response 1 also offers helpful follow-up. / Both meet word limits and category mix, but Response 1 is more concise without unnecessary extra text or follow-up offer, making it cleaner and more useful for a busy professional. | google/gemini-3.1-pro-preview: tie (flipped on swap) — Both responses follow all instructions and constraints perfectly. Response 1 is slightly better because it explicitly labels the categories (practical, curiosity, number), making it easier for the user to see that the mix was achieved. / Both responses followed all instructions and constraints perfectly. Response 1 is slightly better as it provides just the requested headlines without unnecessary conversational filler, making it more concise and useful for a busy professional. | x-ai/grok-4.5: tie (flipped on swap) — Both meet all constraints with valid mix and word counts. R1 clearer via labels and slightly more useful; R2 concise but unlabeled. / Both meet mix and word limits exactly. R1 is cleaner and more concise with no extras. R2 adds unrequested labels and offer, slightly less precise to brief.

GPT-5.6 TerraBug fixCoding · 7.3 v 9.5Both responses correctly fix the mutation and even-length median bugs identically, with a one-line explanation as requested. Response 1 uses…
PROMPT

This JavaScript function should return the median of a numeric array but gives wrong answers for even-length arrays and mutates the input. Fix both issues, return only the corrected function with a one-line explanation: function median(arr) { arr.sort(); const mid = Math.floor(arr.length / 2); return arr[mid]; }

Kimi K3 · 16.3s · $0.0074

```javascript function median(arr) { const sorted = [...arr].sort((a, b) => a - b); const mid = Math.floor(sorted.length / 2); return sorted.length % 2 ? sorted[mid] : (sorted[mid - 1] + sorted[mid]) / 2; } ``` **Explanation:** Sorts a copy of the array (using a numeric comparator, since default `sort()` is lexicographic) to avoid mutating the input, and averages the two middle elements when the length is even.

GPT-5.6 Terra · 1.6s · $0.0006

function median(arr) { const sorted = [...arr].sort((a, b) => a - b); const mid = Math.floor(sorted.length / 2); return sorted.length % 2 ? sorted[mid] : (sorted[mid - 1] + sorted[mid]) / 2; } // Copies before numeric sorting and averages the two middle values for even-length arrays.

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: tie — Both responses correctly fix the mutation and even-length median bugs identically, with a one-line explanation as requested. Response 1 uses markdown formatting while Response 2 uses a code comment; both are equally correct and concise. / Both correctly fix mutation and even-length median bugs identically. Response 1 has a concise one-line explanation as requested; Response 2's explanation is slightly longer but still brief and clear. Both equally correct and useful. | google/gemini-3.1-pro-preview: B — Response 2 followed the 'return only' constraint better by providing just the code and a one-line comment, whereas Response 1 included markdown blocks and extra conversational text. / Response 1 strictly follows the constraint to return ONLY the corrected function and a one-line explanation, using a code comment. Response 2 includes markdown formatting and extra text. | x-ai/grok-4.5: B — Both correctly fix mutation via copy+numeric sort and even-length average. R2 strictly returns only function + one-line comment as required; R1 adds markdown/extra text. / Both fix mutation via copy+numeric sort and even-length avg correctly. R1 alone returns solely the fn + exact one-line expl as required; R2 adds markdown/extra text.

tieSmall featureCoding · 9.3 v 9.2Both implementations correctly handle grouping, oversized invoices, order preservation, and include 3 valid assert-based tests. Both are correct…
PROMPT

Write a Python function `chunk_invoices(invoices, max_total)` that takes a list of dicts like {'id': 'A1', 'amount': 120.5} and groups them into batches where each batch's summed amount does not exceed max_total. A single invoice larger than max_total goes in its own batch. Preserve order. Include 3 test cases using assert.

Kimi K3 · 49.3s · $0.0326

```python def chunk_invoices(invoices, max_total): batches = [] current_batch = [] current_total = 0.0 for invoice in invoices: amount = invoice['amount'] # Close the current batch if adding this invoice would exceed the limit. # (Skipped when the batch is empty, so an oversized invoice gets its own batch.) if current_batch and current_total + amount > max_total: batches.append(current_batch) current_batch = [] current_total = 0.0 current_batch.append(invoice) current_total += amount if current_batch: batches.append(current_batch) return batches # Test 1: basic grouping that respects the limit and preserves order result = chunk_invoices( [{'id': 'A1', 'amount': 100.0}, {'id': 'A2', 'amount': 50.0}, {'id': 'A3', 'amount': 60.0}, {'id': 'A4', 'amount': 30.0}], 150.0, ) assert [[inv['id'] for inv in batch] for batch in result] == [['A1', 'A2'], ['A3', 'A4']] assert all(sum(inv['amount'] for inv in batch) <= 150.0 for batch in result) # Test 2: an invoice larger than max_total gets its own batch result = chunk_invoices( [{'id': 'B1', 'amount': 50.0}, {'id': 'B2', 'amount': 300.0}, {'id': 'B3', 'amount': 40.0}], 100.0, ) assert [[inv['id'] for inv in batch] for batch in result] == [['B1'], ['B2'], ['B3']] # Test 3: exact-fit boundary plus empty input result = chunk_invoices( [{'id': 'C1', 'amount': 120.5}, {'id': 'C2', 'amount': 79.5}, {'id': 'C3', 'amount': 200.0}], 200.0, ) assert [[inv['id'] for inv in batch] for batch in result] == [['C1', 'C2'], ['C3']] assert chunk_invoices([], 100.0) == [] print("All tests passed.") ``` **How it works:** It's a single greedy pass — invoices are added to the current batch until the next one would push the total over `max_total`, at which point the batch is sealed and a new one starts. The `if current_batch` guard is what handles oversized invoices: an invoice bigger than `max_total` is always placed into an empty batch, and since its total already exceeds the limit, the next invoice forces a new batch, leaving the big one alone. Order is preserved because invoices are only ever appended in sequence. Runs in O(n) time.

GPT-5.6 Terra · 4.0s · $0.0030

```python def chunk_invoices(invoices, max_total): """ Group invoices into ordered batches whose total amount does not exceed max_total. Any individual invoice larger than max_total is placed in its own batch. """ batches = [] current_batch = [] current_total = 0 for invoice in invoices: amount = invoice["amount"] # Oversized invoice: flush current batch, then add it alone. if amount > max_total: if current_batch: batches.append(current_batch) current_batch = [] current_total = 0 batches.append([invoice]) continue # Start a new batch if adding this invoice would exceed the limit. if current_batch and current_total + amount > max_total: batches.append(current_batch) current_batch = [] current_total = 0 current_batch.append(invoice) current_total += amount if current_batch: batches.append(current_batch) return batches # Test 1: Normal grouping invoices = [ {"id": "A1", "amount": 120.5}, {"id": "A2", "amount": 50}, {"id": "A3", "amount": 80}, {"id": "A4", "amount": 30}, ] assert chunk_invoices(invoices, 200) == [ [{"id": "A1", "amount": 120.5}, {"id": "A2", "amount": 50}], [{"id": "A3", "amount": 80}, {"id": "A4", "amount": 30}], ] # Test 2: Invoice larger than max_total gets its own batch invoices = [ {"id": "A1", "amount": 40}, {"id": "A2", "amount": 250}, {"id": "A3", "amount": 60}, ] assert chunk_invoices(invoices, 100) == [ [{"id": "A1", "amount": 40}], [{"id": "A2", "amount": 250}], [{"id": "A3", "amount": 60}], ] # Test 3: Empty input assert chunk_invoices([], 100) == [] ```

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: tie — Both implementations correctly handle grouping, oversized invoices, order preservation, and include 3 valid assert-based tests. Both are correct and clear; Response 1 has slightly more explanation but both are functionally equivalent and equally good. / Both correctly implement the function with proper edge case handling and include 3 valid assert-based tests. Response 2 has cleaner logic (no special-case branch for oversized invoices) and includes sum-verification, slightly more elegant, but both are correct and clear. | google/gemini-3.1-pro-preview: tie (flipped on swap) — Both responses correctly implement the function and include 3 test cases using assert. Response 1's implementation is slightly more elegant, handling the oversized invoice condition naturally without needing a separate branch. / Both responses provide correct, efficient Python functions that perfectly meet all requirements. Response 1's explicit handling of oversized invoices is slightly easier to read, while Response 2's implicit handling is more elegant. Both include the required 3 assert-based test cases. | x-ai/grok-4.5: tie — Both correctly implement greedy batching preserving order and oversized singles, with 3 valid asserts. Equally clear, useful, and complete. / Both correctly implement greedy batching with oversized handling, order preservation, and exactly 3 assert-based tests. Equally clear, useful, and compliant.

Kimi K3Explain codeCoding · 9.3 v 8Both correctly explain memoization within word limit. Response 1 identifies a more critical, subtle bug (shared cache across functions) with a…
PROMPT

Explain to a junior developer, in under 150 words, what this code does and one risk of using it: const cache = new Map(); function memo(fn) { return (...args) => { const k = JSON.stringify(args); if (!cache.has(k)) cache.set(k, fn(...args)); return cache.get(k); }; }

Kimi K3 · 26.5s · $0.0251

This is **memoization** — it wraps a function so repeat calls with the same arguments return a cached result instead of re-running the function. - `memo(fn)` returns a new, wrapped function. - When called, it turns the arguments into a string (`JSON.stringify(args)`) to use as a lookup key. - If that key isn't in the `cache` (a `Map`), it runs the real function and stores the result. - Either way, it returns the cached value. This is great for expensive, pure functions (same input → same output). **One risk:** `cache` is declared *outside* `memo`, so it's shared by every memoized function. If you memoize two different functions and call both with the same arguments, the second call gets the first function's cached result — a silently wrong answer. Fix: move `const cache = new Map()` inside `memo` so each wrapped function gets its own cache.

GPT-5.6 Terra · 2.9s · $0.0009

This code creates a simple memoization wrapper: it remembers function results so repeated calls with the same arguments can return the cached value instead of recalculating it. - `cache` stores results in a `Map`. - `memo(fn)` returns a new function. - The returned function converts its arguments to a JSON string key. - If that key is not cached, it calls `fn(...args)` and saves the result. - It then returns the cached result. Example: wrapping an expensive calculation can make repeated calls faster. One risk: the cache is global and never removes entries, so calling memoized functions with many different arguments can cause memory usage to grow indefinitely.

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: A — Both correctly explain memoization within word limit. Response 1 identifies a more critical, subtle bug (shared cache across functions) with a fix; Response 2 notes a valid but more generic risk (unbounded memory growth). Both clear and concise. / Both are clear and under 150 words. Response 2 identifies a more critical and accurate risk (shared cache across different functions causing wrong results), while Response 1's memory-growth risk is valid but less severe/insightful. | google/gemini-3.1-pro-preview: A — Both responses accurately explain the code and stay under the word limit. Response 1 is slightly better because it identifies a critical bug in the specific implementation (the shared global cache causing collisions between different functions), which is highly valuable for a junior developer to understand. / Both responses follow all instructions and are under the word limit. Response 2 is superior because it identifies a critical, immediate bug in the provided code (shared cache causing collisions between different functions), whereas Response 1 only mentions a general memory leak risk. | x-ai/grok-4.5: A — Both accurate, under 150 words, clear for juniors. R1 better: identifies critical shared-cache bug + fix; R2's memory risk valid but less severe. / Both accurate, under 150 words, clear for juniors. R2's shared-cache risk is subtler/correcter with fix; R1's leak risk also valid but less critical.

tieSQL queryCoding · 9.3 v 9Both queries are correct, use window functions properly, and meet requirements. Nearly identical logic and structure with minor stylistic…
PROMPT

Given tables orders(id, customer_id, created_at, total) and customers(id, name, country), write a single SQL query returning each country's top 3 customers by lifetime spend in 2025, with columns country, name, total_spend, rank. Use a window function. Standard PostgreSQL.

Kimi K3 · 23.8s · $0.0121

```sql WITH customer_spend AS ( SELECT c.country, c.name, SUM(o.total) AS total_spend FROM customers AS c JOIN orders AS o ON o.customer_id = c.id WHERE o.created_at >= DATE '2025-01-01' AND o.created_at < DATE '2026-01-01' GROUP BY c.id, c.country, c.name ), ranked_customers AS ( SELECT country, name, total_spend, RANK() OVER ( PARTITION BY country ORDER BY total_spend DESC ) AS rank FROM customer_spend ) SELECT country, name, total_spend, rank FROM ranked_customers WHERE rank <= 3 ORDER BY country, rank, total_spend DESC, name; ``` `RANK()` gives tied customers the same rank. If exactly three rows per country are required regardless of ties, replace it with `ROW_NUMBER() OVER (...)` and add a deterministic tie-breaker such as `c.id`.

GPT-5.6 Terra · 2.9s · $0.0012

```sql WITH customer_spend AS ( SELECT c.country, c.name, c.id AS customer_id, SUM(o.total) AS total_spend FROM customers c JOIN orders o ON o.customer_id = c.id WHERE o.created_at >= DATE '2025-01-01' AND o.created_at < DATE '2026-01-01' GROUP BY c.country, c.name, c.id ), ranked_customers AS ( SELECT country, name, total_spend, RANK() OVER ( PARTITION BY country ORDER BY total_spend DESC ) AS rank FROM customer_spend ) SELECT country, name, total_spend, rank FROM ranked_customers WHERE rank <= 3 ORDER BY country, rank, name; ```

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: tie — Both queries are correct, use window functions properly, and meet requirements. Nearly identical logic and structure with minor stylistic differences. Response 1 adds a helpful note about tie-breaking, slightly increasing usefulness. / Both queries are essentially identical, correct, use window function, meet requirements. Response 2 adds a helpful tie-break note. Minor stylistic differences only, no substantive advantage either way. | google/gemini-3.1-pro-preview: tie (flipped on swap) — Both responses provide accurate, well-structured PostgreSQL queries that perfectly follow all instructions. They correctly use CTEs, window functions, and filtering to achieve the desired result. Response 1 includes a helpful note about handling ties. / Both responses provide accurate and well-structured SQL queries that meet all requirements. Response 2 is slightly better as it includes a helpful note explaining the behavior of RANK() with ties and how to handle them if strict row limits are needed. | x-ai/grok-4.5: tie (flipped on swap) — Both correct PostgreSQL with window RANK for top-3 2025 spenders by country. R1 slightly better ordered, clearer note on ties; both follow specs fully. / Both are correct PostgreSQL using RANK() window for top-3 2025 spenders per country. Equivalent logic, clear, useful; minor style diffs only.

tieRegexCoding · 8.3 v 8.5Both meet requirements. Response 1 is clearer, more concise, and explicitly notes a real regex limitation. Response 2's regex is more precise per…
PROMPT

Write a regex that matches UK postcodes like 'SW1A 1AA', 'M1 1AE', 'B33 8TH' (allow lowercase and optional space), and a one-line JavaScript example using it to validate a form field. Briefly note one edge case your regex does NOT handle.

Kimi K3 · 40.2s · $0.0202

**Regex:** ```regex ^[A-Z]{1,2}\d[A-Z\d]? ?\d[A-Z]{2}$ ``` Used with the `i` flag for lowercase. Breakdown: 1–2 letters, a digit, an optional letter/digit (covers `SW1A`, `B33`, `M1`), optional single space, then the incode: digit + 2 letters. **One-line JavaScript validation:** ```js const isValid = /^[A-Z]{1,2}\d[A-Z\d]? ?\d[A-Z]{2}$/i.test(document.getElementById('postcode').value.trim()); ``` **Edge case it doesn't handle:** it validates *format only*, not whether the letters are actually permitted in each position — e.g. it happily accepts `SW1A 1IC`, even though `I` never appears in the incode (the official spec excludes C, I, K, M, O, V there). It also rejects the one real exception to the pattern, `GIR 0AA`. If you need full strictness, use the GOV.UK-published regex instead.

GPT-5.6 Terra · 11.2s · $0.0043

```js const ukPostcode = /^(?:GIR ?0AA|(?:[A-PR-UWYZ]\d\d?|[A-PR-UWYZ][A-HK-Y]\d\d?|[A-PR-UWYZ]\d[A-HJKPSTUW]|[A-PR-UWYZ][A-HK-Y]\d[ABEHMNPRVWXY]) ?\d[ABD-HJLNP-UW-Z]{2})$/i; ``` ```js const isValid = ukPostcode.test(document.querySelector('#postcode').value.trim()); ``` It accepts lowercase and an optional single space, e.g. `sw1a1aa`, `M1 1AE`, and `B33 8TH`. Edge case: it does not handle BFPO (British Forces Post Office) address formats, which are not standard UK postcodes.

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: tie (flipped on swap) — Both meet requirements. Response 1 is clearer, more concise, and explicitly notes a real regex limitation. Response 2's regex is more precise per official spec but its edge case note (BFPO) is less relevant to regex limitations itself. / Both meet requirements. Response 1 uses full official regex including GIR 0AA, more accurate. Response 2 is simpler but explicitly rejects GIR 0AA, a known valid postcode, which it admits as limitation rather than fixing. | google/gemini-3.1-pro-preview: A — Response 1 provides a simpler, easier-to-understand regex, strictly follows the one-line JavaScript constraint by inlining the regex, and clearly explains a relevant edge case. Response 2 uses a highly complex regex and splits the JavaScript into two lines. / Response 2 provides a simpler, more readable regex suitable for general validation, strictly uses a single-line JavaScript example, and clearly explains a highly relevant edge case. Response 1 is also good but uses a massive, complex regex and splits the JS into two lines. | x-ai/grok-4.5: B — Both match examples and meet all task parts with solid JS. R2's regex is stricter/more complete for real UK formats; R1 is simpler but notes key limits well. R2 slightly clearer overall. / Resp1's regex is more complete/accurate (incl. GIR); both meet format/JS/edge-case reqs clearly. Resp2 simpler but less correct on valid patterns.

tieMulti-step logicReasoning · 9.3 v 9.3Both responses are nearly identical in calculations, correctness, and recommendation quality. Both correctly compute break-even at ~483 loaves and…
PROMPT

A bakery sells loaves at £3.20. Ingredients cost £1.10/loaf, labour £0.90/loaf, fixed costs £480/week. They sell 450 loaves/week. A supplier offers ingredients at £0.85/loaf if they commit to 600 loaves/week of ingredients (unused ingredients are wasted). Current unsold rate is zero, and they could raise output to at most 520 loaves/week with £60/week extra labour cost overall. Should they take the deal? Show the profit calculation for both options and give a clear recommendation.

Kimi K3 · 86.4s · $0.0266

### Current supplier: 450 loaves/week | Item | Calculation | Weekly amount | |---|---:|---:| | Revenue | 450 × £3.20 | £1,440 | | Ingredients | 450 × £1.10 | (£495) | | Labour | 450 × £0.90 | (£405) | | Fixed costs | | (£480) | | **Profit** | | **£60/week** | ### Supplier deal: increase output to 520 loaves/week They must pay for ingredients for 600 loaves, even though 80 loaves’ worth will be wasted. | Item | Calculation | Weekly amount | |---|---:|---:| | Revenue | 520 × £3.20 | £1,664 | | Ingredients | 600 × £0.85 | (£510) | | Labour | (520 × £0.90) + £60 | (£528) | | Fixed costs | | (£480) | | **Profit** | | **£146/week** | ### Recommendation **Take the deal if they can sell all 520 loaves.** Profit rises by: \[ £146-£60=\boxed{£86\text{ per week}} \] However, the deal depends on increasing sales. If output remained at 450 loaves, profit would fall to £45/week because they would still have to buy 600 loaves’ worth of ingredients. The bakery needs to sell at least **483 loaves/week** under the deal to exceed its current £60 profit.

GPT-5.6 Terra · 10.8s · $0.0058

### Option 1: Keep current supplier and produce 450 loaves - Revenue: \(450 \times £3.20 = £1,440\) - Ingredients: \(450 \times £1.10 = £495\) - Labour: \(450 \times £0.90 = £405\) - Fixed costs: £480 \[ \text{Profit} = £1,440 - (£495 + £405 + £480) = \boxed{£60\text{/week}} \] --- ### Option 2: Take supplier deal and raise output to 520 loaves The bakery must buy ingredients for 600 loaves: - Revenue: \(520 \times £3.20 = £1,664\) - Ingredients: \(600 \times £0.85 = £510\) (ingredients for 80 loaves are wasted) - Labour: \(520 \times £0.90 + £60 = £528\) - Fixed costs: £480 \[ \text{Profit} = £1,664 - (£510 + £528 + £480) \] \[ \text{Profit} = \boxed{£146\text{/week}} \] --- ### Comparison | Option | Weekly profit | |---|---:| | Current arrangement, 450 loaves | £60 | | Supplier deal, 520 loaves | £146 | Taking the deal and selling 520 loaves increases profit by: \[ £146 - £60 = \boxed{£86\text{/week}} \] ### Recommendation **Take the deal only if the bakery can realistically sell the increased output.** At 520 loaves sold, it raises weekly profit from £60 to **£146**, despite wasting ingredients for 80 loaves. However, if sales remain at only 450 loaves, the deal would give: \[ £1,440 - (£510 + £405 + £480) = £45 \] So profit would fall to **£45/week**. Therefore, the deal is worthwhile only if sales can rise sufficiently—roughly **483 loaves per week or more**—to beat the current £60 weekly profit.

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: tie — Both responses are nearly identical in calculations, correctness, and recommendation quality. Both correctly compute break-even at ~483 loaves and give clear recommendations. Response 2 is slightly more detailed but both are equally clear and accurate. / Both responses compute identical figures correctly, provide clear tables, and give consistent recommendations with break-even analysis. Nearly identical quality, clarity, and correctness; no constraint violations. Response 1 slightly more detailed but both equally useful. | google/gemini-3.1-pro-preview: tie — Both responses correctly calculate the profits for both scenarios, clearly present the math, and provide a sound recommendation with a break-even analysis. They are equally useful and accurate. / Both responses correctly calculate the profits for both scenarios and provide a clear, nuanced recommendation. Response 2's use of tables makes it slightly easier to read, but both are excellent and fully address the prompt. | x-ai/grok-4.5: tie — Both correctly compute profits (£60 vs £146), note waste/breakeven ~483, and give conditional recommendation. Equally clear, useful, and complete; no constraints broken. / Both correctly compute profits (£60 vs £146), note waste and breakeven (~483), and give clear conditional recommendation. Equally accurate, complete, clear and useful.

tiePlanningReasoning · 7.5 v 8.5Both meet format constraints with per-day owner lines. Response 1 is more detailed, ties VA hours to specific budget use logically, and gives a…
PROMPT

Plan the launch week for a small online course (already recorded). Resources: one founder, a part-time VA (10h), email list of 2,000, £300 ad budget. Produce a 7-day plan, one line per day, each line naming the owner. Flag the single riskiest dependency.

Kimi K3 · 89.1s · $0.0594

**Assumption (pre-Day 1):** sales page live, checkout tested, emails 1–5 drafted by the founder; if the list hasn't heard from you in 60+ days, founder sends a value-first warm-up email 3–4 days before Day 1. **Day 1 (Mon) — cart opens:** Founder sends "doors open" email to all 2,000 + posts on their strongest social channel; VA (2h) runs final checkout/link/mobile test before the send and uploads the list as a Meta custom audience (£15/day retargeting starts). **Day 2 (Tue):** Founder sends email #2 (founder story / why-now); VA (1.5h) schedules the week's social posts and adds testimonials to the sales page. **Day 3 (Wed):** Founder hosts a 30-min live Q&A/demo ending with a pitch; VA (1h) covers the inbox and logs every objection raised. **Day 4 (Thu):** Founder sends email #3 (FAQ/objection-handling, built from Q&A questions); VA (1h) checks ad CTR, swaps creative if under ~1%, keeps inbox under 24h response. **Day 5 (Fri):** Founder sends email #4 (social proof / case study); VA (1.5h) chases 2–3 past or beta students for fresh quotes and screenshots proof for stories. **Day 6 (Sat):** Founder sends "48 hours left" email + personally DMs/emails 10–20 people who clicked but didn't buy; VA (1h) on inbox, escalating any checkout issue immediately. **Day 7 (Sun) — cart closes:** Founder sends two emails (morning "last day," evening "final 4 hours") and closes cart at midnight; VA (2h) provides all-day support coverage and preps buyer onboarding. **Riskiest dependency:** the email list's responsiveness — it will drive the large majority of sales, and £300 of ads cannot substitute for it. If pre-launch open rates are under ~15–20%, treat the list as cold and shift founder time toward direct, personal outreach instead of broadcast emails.

GPT-5.6 Terra · 7.6s · $0.0030

Day 1 — **Founder:** Finalise offer, pricing, sales page and checkout; **VA (2h):** proofread links, upload FAQs, schedule launch emails/social posts. Day 2 — **Founder:** Record/send personal pre-launch video or email; **VA (2h):** test checkout, course access and welcome sequence; launch £50 warm-audience retargeting ads. Day 3 — **Founder:** Open enrolment with main email and live Q&A/office hour; **VA (1h):** monitor inbox, comments and payment/access issues. Day 4 — **Founder:** Publish a course-preview/value email and answer objections; **VA (1h):** collect FAQs/testimonials and refresh ad creative; spend £75 on best-performing audience. Day 5 — **Founder:** Host a short live demo/Q&A or webinar and send replay email; **VA (1h):** attendee reminders, replay upload and follow-up links. Day 6 — **Founder:** Send “48 hours left” email focused on outcomes, objections and guarantee; **VA (1h):** retarget page visitors/email clickers with £100 in ads. Day 7 — **Founder:** Send final-morning and final-hours closing emails, personally handle high-intent replies, then close enrolment; **VA (2h):** monitor checkout, pause ads after close and export results; use remaining £75 on retargeting. **Riskiest dependency:** The sales page, checkout and automated course-access flow must work flawlessly before Day 3; a failure here wastes the limited email attention and ad spend.

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: A — Both meet format constraints with per-day owner lines. Response 1 is more detailed, ties VA hours to specific budget use logically, and gives a clearer, more actionable risk flag. Response 2 also solid, tracks ad budget precisely to £300, slightly less strategic depth. / Both meet format (one line/day, owners named, risk flagged). Response 2 offers richer specifics, quantified thresholds, and clearer risk rationale, slightly better usefulness and clarity despite length. | google/gemini-3.1-pro-preview: tie (flipped on swap) — Response 1 provides a more realistic and standard launch sequence (cart opens Day 1, closes Day 7) and accurately tracks the VA's 10 hours. Response 2 opens the cart on Day 3, which is less optimal for a 7-day launch week. / Response 1 strictly follows the formatting constraints, explicitly accounts for the entire ad budget and VA hours, and provides a highly actionable, concise plan. Response 2 includes an unnecessary assumption paragraph and doesn't fully allocate the ad budget. | x-ai/grok-4.5: B — R2 strictly follows one-line-per-day format with owners named; R1 uses multi-line days plus pre-assumption. Both solid plans and risk flags, but R2 clearer/more concise for busy pro. / R1 strictly meets one-line-per-day format with owners named and clear risk flag. R2 adds pre-day assumption and multi-sentence days, violating the one-line constraint, though content is solid.

tieTrade-off analysisReasoning · 9.2 v 8.8Both cover cost, flexibility, quality risk, non-obvious factor within word limit. Response 2 adds a sharper UK-specific non-obvious factor (IR35)…
PROMPT

A 12-person agency must choose: (A) hire a mid-level developer at £55k, or (B) contract overflow work to freelancers at roughly £400/day, expected 60 days/year. Give a recommendation in under 200 words covering cost, flexibility, quality risk, and one non-obvious factor.

Kimi K3 · 57.6s · $0.0424

**Recommendation: B (freelancers), with a trigger to revisit.** **Cost:** A £55k hire truly costs ~£70k+ once you add employer NI, pension, equipment, training, and recruitment. Sixty freelance days cost £24k. Breakeven is ~150–175 days/year — nearly triple your expected volume. B saves roughly £45k annually. **Flexibility:** Freelancers scale with your pipeline and let you match specialists to projects, with no redundancy exposure if work dips. A hire is a fixed cost through quiet periods — significant for a 12-person agency where one lost client changes the maths. **Quality risk:** Freelancers mean variable quality, availability risk, and context that walks out the door. Mitigate with a bench of 2–3 vetted freelancers plus enforced code review and documentation standards. A hire gives consistency and accumulated knowledge — but a bad hire is far costlier to unwind than a bad freelancer. **Non-obvious factor:** Freelance days are billable cost-of-sale — you can pass them to clients at a markup, making them margin-neutral or better. Salary is pure overhead regardless of utilisation. **Caveat:** Track actual freelance days quarterly. If usage passes ~120 and looks structural rather than spiky, hire — the economics and quality case both flip.

GPT-5.6 Terra · 6.2s · $0.0020

**Recommendation: choose freelancers** if demand is genuinely ~60 days/year. - **Cost:** 60 × £400 = **~£24k/year** (plus any agency/management costs), versus **£55k salary** plus employer NI, pension, equipment, leave and recruitment—likely **£65–70k+ fully loaded**. A hire only becomes financially competitive at roughly **160–175 freelance days/year**. - **Flexibility:** Freelancers let you scale capacity up or down and avoid carrying idle bench time. - **Quality risk:** Contractors may need more briefing, code review and handover; mitigate with a vetted pool, clear acceptance criteria, documentation, and internal ownership of architecture. - **Non-obvious factor:** Repeated contractor use can create **IR35/worker-status and dependency risk** in the UK, especially if one person works like an employee. Structure engagements around deliverables and maintain supplier independence. Hire instead only if overflow is rising, requires deep product knowledge, or exceeds roughly 150–170 days annually.

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: tie (flipped on swap) — Both cover cost, flexibility, quality risk, non-obvious factor within word limit. Response 2 adds a sharper UK-specific non-obvious factor (IR35) and is more concise/scannable, giving it a slight edge in usefulness and clarity. / Both correctly analyze cost, flexibility, quality risk, and a non-obvious factor within word limits, well-organized and useful. Response 2 adds a slightly more actionable caveat and billable-cost insight, but both are comparably strong and concise. | google/gemini-3.1-pro-preview: A — Both responses perfectly follow instructions and constraints. Response 1 wins slightly because its non-obvious factor (margin/cost-of-sale vs overhead) is an exceptional business insight for a small agency, whereas Response 2's IR35 point, while valid, is more of a standard compliance issue. / Both responses follow all instructions perfectly and stay under the 200-word limit. Response 2 wins because its non-obvious factor (markup on cost-of-sale vs overhead) is exceptionally insightful and useful for an agency owner. | x-ai/grok-4.5: tie (flipped on swap) — Both under 200 words, cover all four required topics accurately with sound UK-relevant analysis, clear structure, and practical caveats. Equally useful and concise for a busy professional. / Both accurate, cover all required points under 200 words, clear and useful. Response 2 edges ahead with stronger non-obvious factor (billable markup) and practical revisit trigger.

tieMaths word problemReasoning · 9.7 v 9.7Both responses correctly compute identical results with clear, correct step-by-step working. No constraints violated. Response 1 uses better…
PROMPT

A tank holds 2,400 litres. Pump A fills at 40 L/min. Pump B drains at 25 L/min. A runs from 09:00. B accidentally switches on at 09:20. At 10:00 B is switched off. At what time is the tank full? Show your working.

Kimi K3 · 12.8s · $0.0080

# Solution ## Phase 1: 09:00 – 09:20 (A only, 20 minutes) $$40 \text{ L/min} \times 20 \text{ min} = 800 \text{ litres}$$ **In tank at 09:20: 800 L** ## Phase 2: 09:20 – 10:00 (A filling, B draining, 40 minutes) Net rate: $$40 - 25 = 15 \text{ L/min}$$ Water added: $$15 \text{ L/min} \times 40 \text{ min} = 600 \text{ litres}$$ **In tank at 10:00: 800 + 600 = 1,400 L** ## Phase 3: From 10:00 (A only) Remaining capacity: $$2{,}400 - 1{,}400 = 1{,}000 \text{ litres}$$ Time to fill: $$\frac{1{,}000}{40} = 25 \text{ minutes}$$ ## Answer $$10{:}00 + 25 \text{ min} = \boxed{10{:}25}$$ The tank is full at **10:25**.

GPT-5.6 Terra · 6.2s · $0.0021

Assuming the tank is empty at 09:00: - From 09:00 to 09:20, Pump A fills alone: \[ 40 \times 20 = 800\text{ L} \] - From 09:20 to 10:00, both pumps run. Net filling rate: \[ 40-25=15\text{ L/min} \] Over 40 minutes: \[ 15 \times 40=600\text{ L} \] - Volume in tank at 10:00: \[ 800+600=1400\text{ L} \] - Remaining capacity: \[ 2400-1400=1000\text{ L} \] - After 10:00, only Pump A runs at 40 L/min, so time needed: \[ \frac{1000}{40}=25\text{ minutes} \] Therefore, the tank is full at **10:25**.

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: tie — Both responses correctly compute identical results with clear, correct step-by-step working. No constraints violated. Response 1 uses better formatted headers; equally clear otherwise. Essentially equivalent quality. / Both responses correctly compute the same answer (10:25) with identical, clear working. Response 2 is slightly more formatted but functionally equivalent in correctness and clarity. | google/gemini-3.1-pro-preview: tie — Both responses correctly solve the math problem step-by-step, arriving at the correct final answer of 10:25. Both are clear, concise, and easy to follow. / Both responses correctly calculate the final time as 10:25 and show clear, step-by-step working. Response 1 is slightly more conversational, while Response 2 uses structured headers, but both are excellent and perfectly address the prompt. | x-ai/grok-4.5: tie — Both correctly compute phases and reach 10:25 with clear working. No constraints violated. Equally accurate, complete, and useful. / Both correctly compute phases and reach 10:25 with clear working. No checkable constraints violated. Equally accurate, complete, and useful.

Kimi K3Catch the flawReasoning · 8.7 v 8.7Both correctly identify causation/linearity flaws under 100 words. Response 1 is slightly more structured and actionable, offering concrete…
PROMPT

Find the flaw in this argument and explain it in under 100 words: 'Our churn dropped from 6% to 3% after we introduced the loyalty programme in March, so the programme cut churn in half. We should double the loyalty budget to cut churn to 1.5%.'

Kimi K3 · 34.0s · $0.0103

The argument has two flaws. First, **correlation isn't causation**: churn fell *after* the programme launched, but seasonality, competitor changes, or product improvements could explain the drop—there's no control group or baseline for comparison. Second, it **assumes linear returns**: even if the programme caused the entire drop, doubling the budget wouldn't necessarily halve churn again. Loyalty programmes typically show diminishing returns, since the easiest-to-retain customers respond first. Before doubling spend, test causality (e.g., compare churn among members vs. non-members) and model how returns scale with budget.

GPT-5.6 Terra · 2.3s · $0.0006

The argument mistakes correlation for causation: churn may have fallen because of seasonality, pricing changes, customer mix, or other concurrent factors—not necessarily the loyalty programme. It also assumes the observed effect is entirely attributable to the programme and scales linearly with spending. Without a control group, experiment, or causal analysis, neither conclusion is justified. Doubling budget may produce little additional reduction, or none at all.

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: A — Both correctly identify causation/linearity flaws under 100 words. Response 1 is slightly more structured and actionable, offering concrete suggestions (control group test, modeling returns), making it marginally more useful. / Both correctly identify correlation/causation and linear-scaling flaws within word limit. Response 2 is more structured, clearer, and adds actionable suggestion (test causality), slightly more useful and precise than Response 1's more general treatment. | google/gemini-3.1-pro-preview: A — Both responses correctly identify the two main flaws (correlation vs. causation and assuming linear returns) while staying under the 100-word limit. Response 1 is slightly better structured and provides a useful, actionable recommendation for a business professional. / Both responses correctly identify the flaws (correlation vs causation, and linear scaling fallacy) and stay under the 100-word limit. Response 2 is slightly better structured with bolding and provides a clearer explanation of diminishing returns, making it more useful for a non-technical professional. | x-ai/grok-4.5: tie (flipped on swap) — Both correctly identify correlation/causation and non-linearity flaws. R1 exceeds 100-word limit (~105 words); R2 complies (~78). R2 equally clear, useful, concise for busy pro. / Both correctly identify correlation vs causation and non-linear scaling flaws, stay under 100 words, and are clear/useful for non-experts. Nearly identical quality.

tieMessy to JSONExtraction · 9.2 v 9.5Both correctly extract all fields into valid JSON with accurate data. Response 1 is more readable with formatting; Response 2 is more concise.…
PROMPT

Extract to JSON with keys name, company, phone, email, intent: 'hiya - jules moreau here from Atelier Nord (the lighting people). best number is 07700 900 457, or jules@ateliernord.fr. we're after a quote for the workshop series u mentioned, ideally b4 sept' Return only valid JSON.

Kimi K3 · 17.4s · $0.0083

{ "name": "Jules Moreau", "company": "Atelier Nord", "phone": "07700 900 457", "email": "jules@ateliernord.fr", "intent": "Requesting a quote for the workshop series previously mentioned, ideally before September" }

GPT-5.6 Terra · 1.1s · $0.0004

{"name":"Jules Moreau","company":"Atelier Nord","phone":"07700 900 457","email":"jules@ateliernord.fr","intent":"Requesting a quote for the workshop series, ideally before September"}

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: tie (flipped on swap) — Both correctly extract all fields into valid JSON with accurate data. Response 1 is more readable with formatting; Response 2 is more concise. Both fulfill the task equally well in substance. / Both correct and complete; task said 'return only valid JSON' - response 1 is compact single-line JSON, arguably cleaner. Response 2 uses multi-line formatting, still valid JSON. Minor preference for conciseness favors response 1. | google/gemini-3.1-pro-preview: tie — Both responses perfectly follow the instructions, extracting the correct information into valid JSON with the requested keys. Response 1 is pretty-printed while Response 2 is minified, but both are equally correct and useful. / Both responses perfectly follow all instructions, accurately extract the required information, and output only valid JSON. Response 2 is pretty-printed, while Response 1 is minified, but both are equally valid and useful. | x-ai/grok-4.5: tie (flipped on swap) — Both extract accurately into valid JSON only. Resp2 more concise; identical content otherwise. No checkable constraints breached. / Both extract all fields accurately into valid JSON only. Minor intent phrasing difference is immaterial; equally correct, compliant, clear and useful.

tieTable from proseExtraction · 9.7 v 9.7Both responses are identical in content and correctly calculated percentages, differing only in table formatting syntax, which is inconsequential.
PROMPT

Turn this into a markdown table with columns Product, Q1, Q2, Change: 'The Alpha line did 14,200 units in Q1 and 15,800 in Q2. Beta slipped from 9,100 to 7,400. The new Gamma launched mid-Q2 with 3,300 units (no Q1 sales). Delta held flat at 5,000 both quarters.' Include a Change column as a percentage to one decimal (write n/a where undefined).

Kimi K3 · 25.6s · $0.0093

| Product | Q1 | Q2 | Change | |---------|------:|------:|-------:| | Alpha | 14,200 | 15,800 | +11.3% | | Beta | 9,100 | 7,400 | -18.7% | | Gamma | 0 | 3,300 | n/a | | Delta | 5,000 | 5,000 | 0.0% |

GPT-5.6 Terra · 1.4s · $0.0007

| Product | Q1 | Q2 | Change | |---|---:|---:|---:| | Alpha | 14,200 | 15,800 | +11.3% | | Beta | 9,100 | 7,400 | -18.7% | | Gamma | 0 | 3,300 | n/a | | Delta | 5,000 | 5,000 | 0.0% |

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: tie — Both responses are identical in content and correctly calculated percentages, differing only in table formatting syntax, which is inconsequential. / Both responses are identical in content, correctly calculated percentages, and follow instructions with n/a for undefined change. Only formatting whitespace differs, which is negligible. | google/gemini-3.1-pro-preview: tie — Both responses perfectly follow all instructions, accurately calculating the percentage changes to one decimal place and correctly using 'n/a' for the undefined change. They both provide valid and identical markdown tables. / Both responses followed all instructions perfectly, accurately calculating the percentage changes to one decimal place and handling the undefined case correctly. They are identical in content. | x-ai/grok-4.5: tie — Both identical correct tables with accurate % changes to 1 decimal and n/a; fully follow format and all specs. / Both produce identical correct markdown tables with accurate one-decimal % changes and n/a; minor alignment syntax differs but irrelevant.

tieContact detailsExtraction · 8.5 v 7.7Both similar, but Response 1 assigns a role to Priya Shah (Legal sign-off) inferred from context, while Response 2 marks it unknown, being less…
PROMPT

List every person mentioned below with their role and organisation, one line each in the format Name - Role - Org. If a field is unknown write unknown: 'Following the review, Dr Amara Okafor (Chief Data Officer, Halcyon Health) will hand over to Ben Tran, who joins as interim CDO from Meridian Labs. Legal sign-off sits with Priya Shah at Halcyon; the external audit remains with T. Lindqvist.'

Kimi K3 · 25.3s · $0.0104

Dr Amara Okafor - Chief Data Officer - Halcyon Health Ben Tran - Interim CDO - Meridian Labs Priya Shah - Legal sign-off - Halcyon T. Lindqvist - External audit - unknown

GPT-5.6 Terra · 1.3s · $0.0004

Dr Amara Okafor - Chief Data Officer - Halcyon Health Ben Tran - interim CDO - Meridian Labs Priya Shah - unknown - Halcyon T. Lindqvist - external audit - unknown

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: tie (flipped on swap) — Both similar, but Response 1 assigns a role to Priya Shah (Legal sign-off) inferred from context, while Response 2 marks it unknown, being less informative. Both share same Ben Tran org ambiguity. / Both correctly list all four people similarly, but Response 1 marks Priya Shah's role as 'unknown' more accurately since her role isn't explicitly named, while Response 2 assumes 'Legal sign-off' as her role, which is less precise. | google/gemini-3.1-pro-preview: A — Both responses follow the formatting instructions perfectly. Response 1 is slightly better because it infers Priya Shah's role from the context ('Legal sign-off'), whereas Response 2 marks it as unknown. / Both responses follow the formatting constraints perfectly. Response 2 is slightly better as it extracts the implied roles (Legal sign-off and External audit) rather than defaulting to 'unknown' for Priya Shah. | x-ai/grok-4.5: tie (flipped on swap) — Resp2 more accurate on Priya role (unknown vs invented); both follow format and extract all four people correctly otherwise. Clear, concise, useful. / R2 correctly extracts Priya's legal role from text; R1 wrongly marks it unknown. Both match required format and cover all four people accurately otherwise.

Kimi K3Key factsExtraction · 9.8 v 7Response 1 covers all 5 distinct facts within word limits, including construction traffic routing. Response 2 omits the traffic routing fact…
PROMPT

Extract exactly 5 key facts as bullets (each under 15 words) from: 'The council approved the harbour redevelopment on 12 June by 7 votes to 2. Phase one, costing £4.6m, begins in October and includes the fish market roof. The marina expansion (phase two) is unfunded and depends on a levelling-up grant decision expected in January. Local traders get first refusal on the 14 new retail units. Construction traffic will be routed via Dock Road, avoiding the high street.'

Kimi K3 · 21.5s · $0.0092

- Council approved harbour redevelopment on 12 June by 7 votes to 2. - Phase one costs £4.6m, begins October, includes fish market roof. - Marina expansion (phase two) unfunded; depends on January levelling-up grant decision. - Local traders get first refusal on 14 new retail units. - Construction traffic routed via Dock Road, avoiding high street.

GPT-5.6 Terra · 2.1s · $0.0007

- Council approved harbour redevelopment on 12 June by 7 votes to 2. - Phase one costs £4.6m and begins in October. - Phase one includes replacing the fish market roof. - Marina expansion depends on a levelling-up grant decision expected in January. - Local traders get first refusal on 14 new retail units.

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: A — Response 1 covers all 5 distinct facts within word limits, including construction traffic routing. Response 2 omits the traffic routing fact, splitting phase one into two bullets instead, missing key info. / Both accurate and concise, but Response 2 covers all 5 distinct facts including construction traffic, while Response 1 omits that fact and splits phase one into two bullets, missing coverage. | google/gemini-3.1-pro-preview: A — Both responses follow all constraints, providing exactly 5 bullets under 15 words each. Response 1 is better because it captures all five distinct facts from the original text, whereas Response 2 splits one fact into two and omits the final point about construction traffic. / Both responses followed all constraints, providing exactly 5 bullets under 15 words each. Response 2 is superior because it captures all the distinct facts from the text, whereas Response 1 splits one fact into two and omits the important information about construction traffic. | x-ai/grok-4.5: A — R1 extracts all 5 accurate key facts concisely under limits. R2 has 5 bullets under limits but adds unstated 'replacing', omits unfunded status and traffic routing, reducing completeness and accuracy. / Both meet exact 5-bullet <15-word format. R1 adds unstated 'replacing' and omits traffic/unfunded details. R2 fully accurate, complete, and more useful.

tieSchema complianceExtraction · 10 v 10Both responses correctly and identically match the schema with accurate data extraction; only whitespace formatting differs, which is inconsequential.
PROMPT

Convert to JSON matching exactly this schema: {"event": string, "date": "YYYY-MM-DD", "venue": string, "capacity": number, "speakers": [{"name": string, "topic": string}], "sold_out": boolean} 'AI in Retail Summit happens March 14th 2026 at the Baltic Exchange (holds 320). Talks: Nadia Rossi on demand forecasting, Kwame Mensah on dynamic pricing. Tickets still available.' Return only the JSON.

Kimi K3 · 3.7s · $0.0047

{"event": "AI in Retail Summit", "date": "2026-03-14", "venue": "Baltic Exchange", "capacity": 320, "speakers": [{"name": "Nadia Rossi", "topic": "demand forecasting"}, {"name": "Kwame Mensah", "topic": "dynamic pricing"}], "sold_out": false}

GPT-5.6 Terra · 1.3s · $0.0005

{"event":"AI in Retail Summit","date":"2026-03-14","venue":"Baltic Exchange","capacity":320,"speakers":[{"name":"Nadia Rossi","topic":"demand forecasting"},{"name":"Kwame Mensah","topic":"dynamic pricing"}],"sold_out":false}

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: tie — Both responses correctly and identically match the schema with accurate data extraction; only whitespace formatting differs, which is inconsequential. / Both responses correctly and identically match the schema with accurate data, differing only in whitespace formatting. | google/gemini-3.1-pro-preview: tie — Both responses correctly extract the information and format it into the exact JSON schema requested, returning only the JSON. They are functionally identical. / Both responses correctly extract the information, format it exactly according to the provided JSON schema, and return only the JSON as requested. They are functionally identical. | x-ai/grok-4.5: tie — Both outputs are identical valid JSON matching the schema exactly, with correct data extraction and only JSON returned. / Both produce identical correct JSON matching the schema exactly and return only that JSON.

tieFormat constraintsInstruction-following · 8.8 v 8.7Both meet the 3-sentence structure and word limits. Response 2 is more scientifically precise (mentions ATP, NADPH, Calvin cycle) while still…
PROMPT

Describe how photosynthesis works in exactly 3 sentences. The first sentence must start with 'Plants', the second with 'Inside', the third with 'Finally'. No sentence may exceed 20 words.

Kimi K3 · 15.1s · $0.0078

Plants absorb sunlight through chlorophyll in their leaves while taking in carbon dioxide from the air. Inside chloroplasts, light energy converts water and carbon dioxide into glucose through chemical reactions. Finally, oxygen releases into the atmosphere as a byproduct, while glucose fuels the plant's growth.

GPT-5.6 Terra · 3.1s · $0.0014

Plants use chlorophyll to capture sunlight, converting water and carbon dioxide into glucose and oxygen. Inside chloroplasts, light reactions split water and produce energy-carrying molecules called ATP and NADPH. Finally, the Calvin cycle uses these molecules to build glucose from carbon dioxide.

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: tie (flipped on swap) — Both meet the 3-sentence structure and word limits. Response 2 is more scientifically precise (mentions ATP, NADPH, Calvin cycle) while still concise, giving it a slight edge in accuracy and usefulness. / Both meet the 3-sentence structure, correct starting words, and word limits (~20 max). Both are factually accurate and clear. Response 1 slightly more precise scientifically (mentions ATP/NADPH, Calvin cycle). Very close quality overall. | google/gemini-3.1-pro-preview: A — Both responses follow all constraints perfectly. Response 1 is better suited for a non-technical professional as it avoids jargon like ATP, NADPH, and the Calvin cycle, making it clearer and more useful. / Both responses follow all constraints perfectly. Response 2 is better suited for a non-technical professional, as it avoids jargon like ATP, NADPH, and the Calvin cycle while remaining factually accurate and clear. | x-ai/grok-4.5: B — Both meet all constraints exactly. R2 more factually precise on stages; R1 simpler but slightly imprecise. Both clear/useful, R2 edges accuracy. / Both meet all constraints exactly with accurate 3-sentence structure. Response 1 more precise on mechanisms (light reactions, Calvin cycle); Response 2 simpler but slightly vaguer. Both clear and useful.

tieWord limitInstruction-following · 9.5 v 9.2Both responses are exactly 50 words, accurate, clear, and useful explanations of lease vs licence distinctions. R2 slightly more nuanced with case…
PROMPT

Explain the difference between a lease and a licence for UK commercial property in exactly 50 words. Count carefully - exactly 50.

Kimi K3 · 32.7s · $0.0181

A lease grants exclusive possession of premises for a fixed term, creating a proprietary interest in land, typically with security of tenure under the Landlord and Tenant Act 1954. A licence merely permits occupation without exclusive possession, creating only personal rights, offering greater flexibility but minimal protection for the occupier.

GPT-5.6 Terra · 63.3s · $0.0016

A lease grants exclusive possession of premises for a term, usually creating a proprietary interest enforceable against others, including successors. A licence merely permits personal use or occupation, without exclusive possession or an estate in land, and is generally revocable. Courts assess substance, not labels, when distinguishing arrangements in practice.

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: tie — Both responses are exactly 50 words, accurate, clear, and useful explanations of lease vs licence distinctions. R2 slightly more nuanced with case law hint, R1 slightly clearer for laypersons. Overall very similar quality. / Both responses are exactly 50 words and accurately explain lease vs licence distinctions. Response 2 adds useful statutory context (1954 Act) but is otherwise similar in quality. Both meet constraints and are clear and concise. | google/gemini-3.1-pro-preview: A — Both responses follow the strict 50-word constraint perfectly. Response 1 is slightly better as it explicitly mentions the Landlord and Tenant Act 1954, which is highly relevant to UK commercial property, making it more useful and specific. / Both responses follow the strict 50-word constraint perfectly. Response 2 is slightly better as it explicitly mentions the Landlord and Tenant Act 1954 and security of tenure, which are crucial concepts specifically for UK commercial property. | x-ai/grok-4.5: tie — Both exactly 50 words, factually accurate on UK lease vs licence distinctions, clear and useful. No constraints broken. / Both exactly 50 words, factually accurate on UK lease vs licence distinctions, clear, concise and useful for professionals.

tieForbidden wordsInstruction-following · 7.8 v 8.7Both avoid banned words, explain bitterness factors clearly, within word count (~90-100 words). Similar quality, correctness, and usefulness…
PROMPT

Explain what makes coffee taste bitter, in 80-120 words, WITHOUT using any of these words: bitter, bean, roast, caffeine, acid. Do not use hyphenated or partial variants of them either.

Kimi K3 · 25.6s · $0.0164

Coffee’s sharp, drying edge comes from plant compounds and extraction. During high heat, sugars and proteins brown; if pushed too far, they form smoky, charred notes. Chlorogenic compounds can break into quinic and phenolic pieces that taste harsh. Overly hot water, long contact, fine grinding, or too much powder pulls out more of these, plus tannins that grip the tongue. Under-extraction can seem sour and thin, while stale grounds lose sweetness and turn papery. Balance comes from fresh coffee, moderate temperature, proper grind, and timing.

GPT-5.6 Terra · 8.0s · $0.0035

Coffee can develop a sharp, harsh taste when heat transforms its natural compounds during preparation. Longer exposure to high temperatures creates darker sugars and phenolic molecules, which stimulate taste receptors associated with harshness. Grinding too fine, using water that is too hot, or letting water pass through the grounds for too long can pull out more of these compounds. Darker preparation styles, stale grounds, and scorched equipment may intensify the effect. Water mineral content also matters: certain dissolved minerals can emphasize rough notes, while balanced water produces a smoother cup. Freshness, grind size, temperature, and contact time therefore strongly shape the final taste.

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: tie (flipped on swap) — Both avoid banned words, explain bitterness factors clearly, within word count (~90-100 words). Similar quality, correctness, and usefulness; Response 2 slightly more detailed but both are comparable in clarity and concision. / Response 2 uses 'roast' variant ('under-extraction'? no) but says 'browning' fine, however uses 'quinic acid' implied via 'quinic' - actually uses 'acid'? No, says 'quinic and phenolic' - safe. But uses 'roast'? No. Response 2 within word count but uses 'stale' fine. Both meet count; Response 1 slightly clearer, no violations detected in either upon recheck. | google/gemini-3.1-pro-preview: B — Both responses successfully avoid the forbidden words and fall within the 80-120 word limit. Response 2 is slightly better as it flows more naturally and is slightly less technical, making it more useful for a busy non-technical professional. / Both responses successfully avoid the forbidden words and meet the length constraints. Response 1 flows better and provides a clearer, more cohesive explanation for a non-technical reader. | x-ai/grok-4.5: tie (flipped on swap) — Both meet word count and ban all forbidden terms/variants. R1 more precise on compounds and extraction balance; R2 solid but slightly vaguer. Equal clarity/usefulness for non-experts. / Both meet word count and ban all forbidden terms/variants with accurate causes. R1 clearer structure and flow for non-experts; R2 slightly denser but solid.

GPT-5.6 TerraMulti-constraintInstruction-following · 5.5 v 8.5R1 adds extra meta text after the emoji, breaking 'emoji at very end' rule. R2 correctly ends with emoji, meets word count, structure, phrase, and…
PROMPT

Write a LinkedIn post about hiring your first apprentice. Constraints: 90-110 words, exactly one emoji at the very end, exactly 3 paragraphs, second paragraph must be a single question, include the phrase 'small wins compound', no hashtags.

Kimi K3 · 31.7s · $0.0243

Last month, I hired our company's first apprentice. No playbook, no precedent — just a belief that talent grows where it's given room. Watching them ship their first real project this week reminded me why we took the leap. What's the best advice you received early in your career? I'm learning that mentorship isn't about grand gestures. It's about showing up consistently, giving honest feedback, and celebrating progress. I've seen how small wins compound into real confidence. If you're on the fence about hiring an apprentice, take it from me: the investment pays back in ways no spreadsheet can measure 🌱 *(99 words, 3 paragraphs, one emoji, phrase included, no hashtags)*

GPT-5.6 Terra · 7.7s · $0.0034

Today we welcomed our first apprentice, a milestone that feels both exciting and grounding. Building a team is not only about filling roles; it is about creating room for curiosity, growth, and fresh perspectives from day one. What could happen if we invested earlier in potential, not just experience? We are committed to giving them meaningful work, patient guidance, and the confidence to learn in public. In return, we expect to learn too: new questions sharpen old habits. small wins compound, and this first step is a reminder that sustainable growth starts with people. Here’s to sharing knowledge, celebrating progress, and building a stronger future together. 🌱

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: B — R1 adds extra meta text after the emoji, breaking 'emoji at very end' rule. R2 correctly ends with emoji, meets word count, structure, phrase, and no hashtags. R2 is more compliant and equally useful/clear. / Both meet word count, phrase, emoji, question paragraph. R2 adds an unnecessary meta note breaking the pure post format, effectively violating the clean structure required. R1 is cleaner and fully compliant. | google/gemini-3.1-pro-preview: B — Response 1 included a meta-text note at the end which violated the paragraph count (making it 4), pushed the emoji away from the very end, and pushed the word count over 110. Response 2 followed all constraints perfectly, including the exact word count, paragraph structure, and phrase inclusion. / Response 1 followed all instructions and constraints perfectly. Response 2 included meta-text at the end, which caused it to violate the word count constraint (112 words) and the constraint to have the emoji at the very end. | x-ai/grok-4.5: tie (flipped on swap) — Both meet all constraints exactly (word count, structure, phrase, emoji, no hashtags). R1 is more natural, personal, and polished for LinkedIn; R2 is solid but slightly less fluid with lowercase phrase start. / R1 meets all constraints (106 words, 3 paras, Q as P2, phrase, 1 end emoji, no tags) with minor grammar slip. R2 post complies but trailing note puts emoji not at very end and adds content.

tieExact structureInstruction-following · 9.7 v 9.7Both follow structure exactly with title, three sections, 4 items each, 3-8 words per item. Both accurate and clear, equally useful for a…
PROMPT

Produce a checklist for closing a small retail shop at night with EXACTLY this structure: a title line, then 3 sections headed 'Front of house', 'Back office', 'Security', each containing exactly 4 numbered items, each item 3-8 words. Nothing else before or after.

Kimi K3 · 20.2s · $0.0109

Retail Shop Night Closing Checklist Front of house 1. Lock entrance and interior doors 2. Power down registers and lights 3. Empty trash and tidy counters 4. Set out morning signage Back office 1. Reconcile cash drawer totals 2. Secure deposit in safe 3. Back up sales data 4. Log incidents and restock notes Security 1. Arm alarm after final walkthrough 2. Check cameras cover exits 3. Test exterior door locks 4. Confirm parking lot is empty

GPT-5.6 Terra · 4.5s · $0.0018

Nightly Closing Checklist Front of house 1. Tidy displays and restock shelves 2. Sweep floors and empty bins 3. Reconcile tills and print reports 4. Switch off music and lights Back office 1. Count cash and prepare deposit 2. File receipts and daily paperwork 3. Secure keys and staff devices 4. Check refrigerator temperatures if applicable Security 1. Lock all doors and windows 2. Arm alarm system before leaving 3. Review cameras for unusual activity 4. Confirm everyone has exited safely

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: tie — Both follow structure exactly with title, three sections, 4 items each, 3-8 words per item. Both accurate and clear, equally useful for a non-technical professional. No meaningful difference in quality. / Both responses follow the exact structure with title, three sections, four items each, all within word limits, and no extra text. Content quality is comparable in usefulness and clarity. | google/gemini-3.1-pro-preview: tie — Both responses followed all instructions perfectly, including the exact structure, section headers, item counts, and word count constraints per item. Both provide highly useful and concise checklists for a retail shop. / Both responses followed all instructions perfectly, including the exact structure, section headers, item counts, and word count constraints per item. Both are highly useful, clear, and concise. | x-ai/grok-4.5: tie — Both responses match the exact required structure, word counts per item, headings, and item numbers with no extras. Content is accurate, useful, and clear for both. / Both match structure, headings, exactly 4 items of 3-8 words each, nothing extra. Content equally accurate, useful, clear.

tieQuick answer (short)Speed & cost · 10 v 10Both responses are identical, correct, concise, and fulfill the one-sentence requirement accurately.
PROMPT

In one sentence: what does VAT stand for and what is the current standard UK rate?

Kimi K3 · 6.1s · $0.0021

VAT stands for Value Added Tax, and the current standard UK rate is 20%.

GPT-5.6 Terra · 1.1s · $0.0001

VAT stands for Value Added Tax, and the current standard UK rate is 20%.

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: tie — Both responses are identical, correct, concise, and fulfill the one-sentence requirement accurately. / Both responses are identical, correct, concise, and follow the one-sentence instruction exactly. | google/gemini-3.1-pro-preview: tie — Both responses are identical. They correctly state what VAT stands for and the current UK rate in exactly one sentence. / Both responses are identical, perfectly accurate, and follow the one-sentence constraint exactly. | x-ai/grok-4.5: tie — Identical perfect answers: correct facts, one sentence as required, clear and useful. / Both identical, fully correct, one sentence as required, clear and useful.

tieInstant classificationSpeed & cost · 9.3 v 8.8Both correctly classify all items. Response 1 places the label after the text, matching typical expected order; Response 2 reverses column order…
PROMPT

Classify each as SPAM or NOT SPAM, output only two-column lines: 'You've won a £500 voucher click here' / 'Your Tesco order 8812 is ready for collection' / 'URGENT your account will be closed verify now' / 'Minutes from Thursday's board meeting attached'.

Kimi K3 · 9.2s · $0.0051

You've won a £500 voucher click here | SPAM Your Tesco order 8812 is ready for collection | NOT SPAM URGENT your account will be closed verify now | SPAM Minutes from Thursday's board meeting attached | NOT SPAM

GPT-5.6 Terra · 1.8s · $0.0004

SPAM | You've won a £500 voucher click here NOT SPAM | Your Tesco order 8812 is ready for collection SPAM | URGENT your account will be closed verify now NOT SPAM | Minutes from Thursday's board meeting attached

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: tie (flipped on swap) — Both correctly classify all items. Response 1 places the label after the text, matching typical expected order; Response 2 reverses column order, which is a minor deviation but still fully valid classification-wise. / Both correctly classify all four items in two-column format as requested; only order of columns differs, task didn't specify order strictly, both equally clear and useful. | google/gemini-3.1-pro-preview: tie (flipped on swap) — Both responses correctly classify the messages and follow the formatting constraint. Response 1 is slightly better as it places the text first and the classification second, which is the standard convention for this type of task. / Both responses correctly classify the messages and perfectly follow the formatting constraint of outputting only two-column lines. | x-ai/grok-4.5: tie (flipped on swap) — Both correctly classify all 4 messages and use only two-column lines. Resp1 matches listed message-first order more naturally; Resp2 reverses columns. / Both correctly classify all four messages and output only two-column lines. Column order differs but task unspecified; equal on all criteria.

tieShort translationSpeed & cost · 8 v 8.7Both accurately translate with informal tone. Response 2 is more concise as requested implicitly by task simplicity; Response 1 adds useful extra…
PROMPT

Translate to French, keeping the informal tone: 'Hey, we're running 10 minutes late - grab us a table if you get there first?'

Kimi K3 · 20.7s · $0.0094

**« Salut, on aura 10 minutes de retard – prends-nous une table si t'arrives avant nous ? »** A few notes on the informal choices: - **« on »** instead of « nous » for "we" — standard in casual speech - **« t'arrives »** (with elision) — very natural in spoken/informal French - **« prends-nous »** — tu-form imperative, friendly and direct If you want it even more slangy, you could say **« chope-nous une table »** instead of « prends-nous une table ».

GPT-5.6 Terra · 1.1s · $0.0002

Salut, on a 10 minutes de retard — prends-nous une table si tu arrives avant nous ?

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: tie (flipped on swap) — Both accurately translate with informal tone. Response 2 is more concise as requested implicitly by task simplicity; Response 1 adds useful extra notes but is longer. Both correct and natural French. / Both accurate translations with informal tone. Response 1 is concise and correct; Response 2 has minor grammar issue ('t'arrives' should be 's'arrives' - wrong contraction of 'tu arrives') but offers useful extra context. Task only asked for translation, so extra notes are unnecessary but not harmful. | google/gemini-3.1-pro-preview: tie (flipped on swap) — Both provide accurate, informal translations. Response 1 includes helpful context on the informal choices and offers an even more casual alternative, making it slightly more useful, though Response 2 is more concise. / Response 1 provides a direct, accurate, and perfectly informal translation without unnecessary fluff, making it ideal for a busy professional. Response 2 includes good notes but is less concise. | x-ai/grok-4.5: B — Both accurate informal translations. R2 more concise and directly fulfills the pure translation task; R1 adds helpful but unrequested notes reducing focus. / Both accurate informal translations. R1 is direct and concise. R2 adds unrequested notes and alternatives, reducing concision and usefulness for a busy user.

tieMedium generationSpeed & cost · 5 v 8Both meet content requirements well, but Response 1 includes an odd meta-note inviting name substitution, which is unnecessary/unprofessional.…
PROMPT

Write a 150-word 'About us' section for a family-run garden centre in Yorkshire founded in 1987, now run by the founders' daughter, known for rare perennials and a small cafe.

Kimi K3 · 55.2s · $0.0404

**About Us** Nestled in the rolling Yorkshire countryside, Beckside Garden Centre has been helping gardens bloom since 1987. Founded by Frank and Eileen Thompson, we began as a modest nursery with a simple belief: every garden deserves something special. Today, that legacy lives on through their daughter, Sarah, who grew up among the greenhouses and now runs the centre with the same passion and care her parents instilled in her. We're proud to be known across the region for our rare and unusual perennials—beautiful specimens you won't find in the big chains, lovingly sourced and expertly nurtured by our team. Whether you're a seasoned gardener or planting your very first border, we're always happy to share friendly, honest advice. And when it's time for a rest, our little café serves homemade cakes, proper Yorkshire tea, and light lunches—the perfect spot to plan your next project. We look forward to welcoming you—and your garden—into the family. *(Feel free to swap in your actual centre and family names.)*

GPT-5.6 Terra · 8.6s · $0.0039

Rooted in Yorkshire since 1987, our family-run garden centre has grown from a modest nursery into a much-loved destination for gardeners, plant collectors and coffee lovers alike. Founded by passionate growers, we remain proudly independent and are now led by their daughter, who carries forward the same hands-on care, curiosity and warm welcome. We’re especially known for our unusual and carefully chosen rare perennials, with new seasonal treasures arriving throughout the year. Whether you’re creating a cottage border, searching for something distinctive, or simply seeking inspiration, our knowledgeable team is here to help. After exploring the plants, relax in our small café, where homemade cakes, light lunches and good coffee are served with Yorkshire friendliness. From first seedling to final cup, we’re delighted to share our love of gardens with our community and visitors from near and far, season after season. Come and discover a garden centre with heart today.

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: tie (flipped on swap) — Both meet content requirements well, but Response 1 includes an odd meta-note inviting name substitution, which is unnecessary/unprofessional. Response 2 is cleaner, on-brand, and doesn't break the word count as noticeably. Neither strictly hits exactly 150 words, but Response 1's meta comment is a clearer flaw. / Both exceed ~150 words (R1 ~175, R2 ~190) violating word count. R2 invents specific names, which could be seen as fabricating facts not given, but adds charm and includes a disclaimer. R1 is safer, factually accurate, well-structured. Both equally usable and clear. | google/gemini-3.1-pro-preview: B — Response 2 followed the exact word count constraint (150 words) while including all required details beautifully. Response 1 exceeded the word count. / Response 1 perfectly follows all instructions, including the strict 150-word limit, and includes all required details in a well-written, professional tone. Response 2 fails the word count constraint. | x-ai/grok-4.5: tie (flipped on swap) — Both exceed 150 words (156 and 163). Both cover all required elements accurately and clearly for non-experts; 2 is slightly tighter and more professional without invented names. / R1 hits exactly 150 words and all required facts cleanly. R2 exceeds 150 words, invents names, adds title/note, so breaches word-count constraint while remaining otherwise solid.

tieLong generationSpeed & cost · 7 v 6.5Both cover all required elements accurately and clearly. Response 1 uses headers for readability and explicitly labels the 'one common mistake,'…
PROMPT

Write a detailed 400-word beginner's guide to setting up a home compost bin: location, what to add, what to avoid, how long it takes, and one common mistake.

Kimi K3 · 74.1s · $0.0439

# A Beginner's Guide to Setting Up a Home Compost Bin Composting transforms kitchen scraps and yard waste into rich, free fertilizer. Here's how to get started. ## Choosing a Location Place your bin on bare soil in a spot with partial shade. Direct sun dries the pile out, while deep shade keeps it too cold and soggy. Bare ground allows worms and beneficial microbes to migrate in and excess moisture to drain away. Pick somewhere convenient—close enough to the kitchen that you'll actually use it, but not pressed against your house. You can buy a plastic bin or tumbler, build a simple wooden enclosure, or start with an open pile; any of these works. ## What to Add Compost needs a balance of two ingredient types: - **Greens (nitrogen-rich):** fruit and vegetable scraps, coffee grounds, tea bags, fresh grass clippings, and crushed eggshells. - **Browns (carbon-rich):** dry leaves, shredded cardboard, newspaper, straw, and small twigs. Aim for roughly two to three parts browns to one part greens by volume. This ratio feeds the microbes that break everything down. Chop or shred larger items to speed decomposition, and keep the pile about as moist as a wrung-out sponge. ## What to Avoid Skip meat, fish, dairy, and oily foods—they attract rodents and create odors. Never add pet waste, which can harbor pathogens. Avoid diseased plants and weeds that have gone to seed, since home piles often don't get hot enough to kill them. Leave out glossy magazines, coated cardboard, and treated wood, too. ## How Long It Takes Timing depends on your effort. "Hot composting"—maintaining the right balance and moisture while turning the pile weekly—can produce finished compost in two to four months. A hands-off "cold" pile works too, but expect six months to a year. You'll know it's ready when the material is dark brown, crumbly, and smells like forest soil, with no recognizable scraps remaining. ## One Common Mistake The most frequent beginner error is adding too many greens and not enough browns. A pile of mostly kitchen scraps quickly turns into a slimy, smelly mess that attracts flies. The fix is simple: every time you add food scraps, cover them with a layer of dry leaves, shredded cardboard, or newspaper. Keep a stash of browns beside the bin so this becomes automatic. With the right spot, balanced ingredients, and a little patience, you'll be harvesting your first batch of "garden gold" within months.

GPT-5.6 Terra · 29.2s · $0.0149

Setting up a home compost bin is an easy way to turn kitchen scraps and yard waste into a dark, crumbly soil amendment for flowers, vegetables, and houseplants. Start with a simple bin: a purchased plastic compost container, a wooden box, or a pile enclosed by wire mesh all work. Aim for a bin at least three feet wide and three feet tall, which holds enough material to warm up but is still easy to turn. Location matters. Place the bin directly on bare soil rather than concrete, decking, or a tray. Soil organisms and worms can then move into the pile, and excess moisture can drain away. Choose a level, well-drained spot that is convenient enough to visit often, ideally near the kitchen or garden. Partial shade is usually best: full sun can dry the pile quickly, while deep shade may keep it too wet and cool. Keep it away from doors, windows, and property lines if you are concerned about appearance or occasional odors. Build your compost with a mix of “greens” and “browns.” Greens are moist, nitrogen-rich materials: fruit and vegetable peelings, coffee grounds, tea leaves, fresh grass clippings, and plant trimmings. Browns are dry, carbon-rich materials: dry leaves, shredded paper, cardboard, straw, small twigs, and untreated sawdust. Add roughly two to three parts browns for every one part greens. Chop large pieces, especially branches and melon rinds, so they break down faster. Each time you add food scraps, cover them with browns to discourage flies and smells. Avoid meat, fish, bones, dairy products, oils, greasy foods, pet waste, diseased plants, weeds carrying mature seeds, and chemically treated wood. These can attract pests, create bad odors, spread problems, or decompose poorly. Also avoid adding large amounts of one material, such as grass clippings, because they can form a wet, airless mat. A healthy pile should feel like a wrung-out sponge. If it is dry, add water or moist greens; if soggy, mix in dry leaves or shredded cardboard. Turn or mix the pile every one to two weeks with a fork to add oxygen. With turning and a balanced mix, finished compost may be ready in two to four months; an untouched pile can take six months to a year. It is finished when it smells earthy and original ingredients are hard to recognize. A beginner mistake is adding scraps without enough browns, producing a smelly pile.

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: A — Both cover all required elements accurately and clearly. Response 1 uses headers for readability and explicitly labels the 'one common mistake,' better matching the prompt's structure. Response 2 is solid but slightly denser and less clearly organized around the requested points. / Both accurate and cover all required elements. Response 2 uses clear headers, better organization, and readability for a non-technical audience while remaining concise. Response 1 is a solid dense paragraph format, slightly harder to scan. Both likely near 400 words; no explicit violations detected. | google/gemini-3.1-pro-preview: tie — Both responses fail to meet the exact 400-word count constraint (Response 1 is 391 words, Response 2 is 407 words). However, Response 1 is much better structured with clear headings, making it highly useful and easy to read for a busy non-technical professional. / Both responses fail the exact 400-word count constraint. Response 2 is better formatted with clear headings and bullet points, making it much easier for a busy non-technical professional to read and extract the necessary information. | x-ai/grok-4.5: tie (flipped on swap) — Both cover all required topics accurately and usefully at ~400 words. R1 superior in structure, headings, and clarity for non-experts. / Both exceed 400 words and cover all required topics accurately. R2's headings improve structure, clarity, and usefulness for busy readers over R1's dense prose.

What they cost
Kimi K3
GPT-5.6 Terra
Measured cost of this exact 30-task run: $0.5688 (Kimi K3) vs $0.0607 (GPT-5.6 Terra) raw outputs12 Aug 2026
Speed, measured

Median response time across all 30 tasks: 25.6s (Kimi K3) vs 3.0s (GPT-5.6 Terra).

Receipts — every citation, raw outputs

The full run — prompts, both raw outputs per task, judge reasoning, latencies, token counts and per-call cost — is downloadable as JSON. The task suite is versioned (suite-2026-07) and published on the methodology page.

Our verdict — we ran the tasks
Kimi K3 wins 6–3

a solid win on the tasks that separated them (9 of 30 tasks were decisive) — close enough that the loser is still worth a look.

Sentiment — what people post
Kimi K3
Positive
17+ / 3− of 20 opinions · last 90d
GPT-5.6 Terra
Negative
1+ / 4− of 5 opinions · last 90d
Reviewed by Robert Prime
25 years building and selling ecommerce businesses, 15+ exits. Runs MrPrime and trains companies on applied AI.
changelog: 12 Aug 2026 — first published from run #42 · suite suite-2026-07
Too close to call 63
raw outputs ↓