Battles / Best free

Probably Gemini 3.5 Flash

Gemini 3.5 Flash vs DeepSeek V4 Flash · Best free

Gemini 3.5 Flash came out in front, but not by enough for us to call it proven on a suite this size. Treat it as the way to bet, not as a settled result.

92
of 11 decided · 19 tied

Gemini 3.5 Flash took 9 of the 11 tasks that had a clear winner (Gemini 3.5 Flash 9, DeepSeek V4 Flash 2). The judge could pick a winner on 11 of 30 tasks; on the other 19 it could not tell them apart. That is a lean, not a proven win — at this sample size we cannot rule out chance, so we are not calling it decisive.

So which should you pick?

Worth knowing: DeepSeek V4 Flash costs noticeably less.

DeepSeek V4 Flash was 110.0x cheaper to run - $0.0036 against $0.3960 across the whole suite. Treat that with caution: DeepSeek V4 Flash's API price moved 59% in the last 7 days ($0.0882-$0.14 per million input tokens), because the provider blends it across whoever is serving the model
tested 7 Aug 202630 tasks3 judges from different labs
Show the full workings
last verified 7 Aug 2026previously Gemini 3.5 Flash 61 DeepSeek V4 Flash, published 7 Aug 2026 under judge-2026-08 — replaced: re-run under judge-2026-08suite suite-2026-07judges: anthropic/claude-sonnet-5, openai/gpt-5.6-terra, x-ai/grok-4.5 (none of them a contestant)judge protocol judge-2026-08 — constraint cap enforced in code, applied to 42 judge scores in this runnot statistically decisive — Wilcoxon signed-rank on score margins p=0.0147 (n=26); sign test on win counts p=0.0654 — held to our confidence gatejudge swap-consistency (Cohen’s κ) = 0.507moderate agreement κ measures how often a judge reached the same verdict when the two answers were shown in the opposite order, corrected for agreement you would get by chance. A battle below 0.21 does not publish a winner at all; above it, this number is how much of the scoreline survived the swap test rather than a claim that the judges were reliable.

The evidence

Suite-by-suite
Writing
230
Coding
230
Reasoning
140
Extraction
230
Instruction-following
221
Speed & cost
041

blue = Gemini 3.5 Flash wins · grey = ties · white = DeepSeek V4 Flash wins (5 tasks per suite)

Round-by-round — all 30 tasks
Gemini 3.5 FlashCold emailWriting · 8.2 v 7.2Both meet word limit, tone, and CTA requirements. Response 1 is more personable and specific (team size, favourite beer), fitting the 6-person…
PROMPT

Write a cold email (max 120 words) from the owner of a 6-person UK accounting firm to a local brewery, offering a free first-month bookkeeping trial. Friendly but professional, one clear call to action, no buzzwords.

Gemini 3.5 Flash · 9.0s · $0.0124

Subject: Bookkeeping for [Brewery Name] (from a local fan) Hi [Name], I’m [My Name], and I run [Firm Name], a team of six accountants based just down the road in [Town]. I'm a big fan of your [insert favourite beer, e.g., pale ale]. Between managing brews, suppliers, and beer duty, I know how easily the paperwork can pile up. We help local independent businesses keep their numbers sorted so they can focus on brewing. To show you how we can take the weight off your shoulders, I’d love to handle your bookkeeping for free for your first month—no strings attached. Do you have 10 minutes for a quick phone call next Tuesday to see if we're a good fit? Best regards, [My Name] [Firm Name]

DeepSeek V4 Flash · 6.7s · $0.0001

Subject: A practical idea for [Brewery Name]’s books Hi [Name], I run a small accounting firm just down the road. I’ve been watching your steady growth and know that for breweries, the bookkeeping can get tangled—especially with stock, duty, and margins. Here’s a simple offer: I’ll handle your first month’s bookkeeping, free. No commitment needed. You’ll see exactly how we work and whether it fits. If it does, great; if not, no hard feelings. I’d love to offer you something practical upfront. If you’re interested, just reply and I’ll set up a quick call. Best, [Your Name] [Your Firm]

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: A — Both meet word limit, tone, and CTA requirements. Response 1 is more personable and specific (team size, favourite beer), fitting the 6-person firm detail. Response 2 is solid but slightly less warm and omits firm size detail. / Both meet constraints and word limit, friendly/professional tone, one CTA, no buzzwords. Response 2 has slightly more specific personalization and clearer CTA (specific day/time), making it marginally more effective for a busy reader. | openai/gpt-5.6-terra: tie (flipped on swap) — Response 1 clearly identifies a six-person local firm, offers a free first month, uses brewery-specific context, and has one direct call to action. Response 2 is concise and professional but omits the required six-person detail and does not establish the firm as UK-based. / Response 1 is concise, friendly, and has one clear CTA, but does not state the firm has six people or is UK-based. Response 2 includes the six-person detail and a clear CTA, but exceeds 120 words (about 125), omits explicit UK context, and includes an awkward beer placeholder. | x-ai/grok-4.5: A — Both under 120 words, no buzzwords, clear CTA. R1 better meets all specs (mentions 6-person firm) and is more personalized/useful. / Both under 120 words, professional, one CTA, no buzzwords. Resp2 correctly notes 6-person firm and is more tailored/useful; Resp1 omits size detail.

Gemini 3.5 FlashProduct descriptionWriting · 8.8 v 5.7Both are factually accurate, avoid cliches, and target gym-goers well. R1 meets the 60-80 word range (67 words) while R2 falls short at 56 words…
PROMPT

Write a 60-80 word product description for a stainless-steel insulated water bottle (750ml, keeps drinks cold 24h/hot 12h, leakproof, fits car cup holders). Target audience: gym-goers. Avoid cliches like 'stay hydrated in style'.

Gemini 3.5 Flash · 9.4s · $0.0146

Power through your toughest workouts with this heavy-duty 750ml stainless-steel bottle. Built for high-performance training, it keeps your water ice-cold for 24 hours or your pre-workout hot for 12. Toss it in your gym bag without a second thought—the leakproof seal guarantees zero spills. With a slim profile that fits standard car cup holders, it’s the ultimate companion for your daily commute and every rep in between.

DeepSeek V4 Flash · 2.1s · $0.0000

Crush your workout without stopping for refills. This 750ml stainless-steel bottle keeps ice water cold for 24 hours or post-workout coffee hot for 12. The leakproof lid survives a tossed gym bag, while the tapered base fits most car cup holders for easy transport. Durable, sweat-free, and built to handle heavy reps—just fill, grip, and go.

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: tie (flipped on swap) — Both are factually accurate, avoid cliches, and target gym-goers well. R1 meets the 60-80 word range (67 words) while R2 falls short at 56 words, violating the explicit word count constraint. R1 is slightly more detailed and complete. / Both meet word count and avoid clichés. Response 1 (63 words) is tighter and more vivid; Response 2 (72 words) uses 'ultimate companion,' a mild cliché, slightly weaker. Both accurate and on-brief; Response 1 edges out on concision and freshness. | openai/gpt-5.6-terra: A — Response 1 is 66 words, covers every specification, targets gym-goers, and avoids the prohibited cliché. Response 2 is engaging and covers the core features, but it is only 55 words, below the required 60–80 range. Its unsupported “sweat-free” claim is also less precise. / Response 1 is accurate, gym-focused, and clear, but contains 56 words, below the required 60–80 range. Response 2 meets the word count at 68 words and covers all features clearly. Both use effective, non-technical language; Response 2 is slightly more promotional and includes a mild overclaim (“guarantees zero spills”). | x-ai/grok-4.5: A — Resp1 hits 68 words, covers all specs accurately for gym audience without cliches. Resp2 is only 57 words (under limit), otherwise solid but incomplete on length constraint. / Both cover specs accurately for gym audience without cliches and are clear/useful. R1 has 57 words; R2 has 68 and slightly stronger flow.

tieSummarise messy notesWriting · 8.5 v 9.2Both accurate and concise. Response 2 includes 'next meeting Tuesday' detail, omitted in Response 1, making it slightly more complete while…
PROMPT

Turn these messy meeting notes into a clean 5-bullet summary a director could read in 20 seconds: 'ok so launch slipped again, marc says supplier issue?? maybe 3 wks. sarah wants to push the email campaign anyway but tom thinks wait. budget - we're 8k over, mostly the packaging redo. Q: do we tell retail partners now or after new date confirmed. also NEED to hire the warehouse temp before august rush. next mtg tues.'

Gemini 3.5 Flash · 5.4s · $0.0073

Here is a 20-second summary for a director: * **Launch Delayed:** The launch is pushed back by approximately 3 weeks due to supplier issues. * **Budget Overrun:** We are currently $8k over budget, driven primarily by the packaging redesign. * **Marketing Alignment:** The team is divided on whether to launch the email campaign now or pause it until the new date is set. * **Pending Decision:** We need to decide whether to notify retail partners of the delay immediately or wait for a confirmed launch date. * **Urgent Hiring:** A warehouse temp must be hired before the August peak season.

DeepSeek V4 Flash · 3.3s · $0.0001

Here’s a clean 5‑bullet summary for a director: - **Launch slipped ~3 weeks** due to supplier issue (Marc). - **Email campaign timing** split: Sarah wants to push forward, Tom advises holding. - **Budget overrun $8k** – primarily from packaging redo. - **Retail partner communication** – decide now or wait until new date is confirmed. - **Hire warehouse temp** before August rush. Next meeting Tuesday.

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: tie (flipped on swap) — Both accurate and concise. Response 2 includes 'next meeting Tuesday' detail, omitted in Response 1, making it slightly more complete while staying within 5 bullets. / Both accurately capture all points in 5 clean bullets, concise and director-ready. Response 1 includes next meeting note, slightly more complete; Response 2 slightly more polished wording. Essentially equal quality. | openai/gpt-5.6-terra: B — Both provide exactly five clear, accurate bullets. Response 1 is polished but omits the next meeting date and presents the supplier cause more definitively than the notes. Response 2 captures every key point, including ownership of the campaign disagreement and Tuesday’s meeting, while remaining more concise. / Both are accurate, concise, and use exactly five bullets. Response 1 captures every material note, including the Tuesday meeting, while Response 2 omits it. Response 2 is slightly more polished in phrasing, but Response 1 is more complete and equally director-friendly. | x-ai/grok-4.5: tie (flipped on swap) — Both deliver accurate 5-bullet summaries covering all key points clearly and concisely for a director. No constraints breached. / Both accurate 5-bullet summaries. R1 includes all key points incl. next mtg; R2 omits it. Equal clarity/usefulness.

tieTone rewriteWriting · 7.8 v 9Both accurately preserve facts, are warm, take responsibility, and are under 90 words. Response 2 is slightly more concise; Response 1 slightly…
PROMPT

Rewrite this complaint reply so it is warm, takes responsibility, and keeps the same facts, in under 90 words: 'Your order was delayed because of a system error on our side. The refund has been processed and will arrive in 3-5 days. We cannot offer further compensation. Let us know if you have questions.'

Gemini 3.5 Flash · 5.1s · $0.0080

I am so sorry your order was delayed. This was due to a system error on our end, and we take full responsibility for the frustration this caused. Your refund has been processed and will arrive in 3-5 business days. While we unfortunately cannot offer further compensation, we truly appreciate your patience and understanding. Please let us know if you have any questions—we are always here to help.

DeepSeek V4 Flash · 4.7s · $0.0001

We're truly sorry your order was delayed due to a system error on our end. We've processed your refund, which should arrive in 3–5 days. While we can't offer additional compensation, we hope this resolves the issue. Please let us know if you have any further questions—we're here to help.

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: tie — Both accurately preserve facts, are warm, take responsibility, and are under 90 words. Response 2 is slightly more concise; Response 1 slightly warmer. Essentially equivalent quality. / Both responses accurately preserve facts, take responsibility, stay under 90 words, and are warm and clear. R2 is slightly more empathetic with added phrasing, but both are equally effective and concise. | openai/gpt-5.6-terra: B — Both are warm, clear, under 90 words, and retain the key facts. Response 1 more explicitly accepts responsibility but changes “3–5 days” to “3–5 business days,” adding an unsupported detail. Response 2 preserves the timeline exactly and is more concise, though its ownership language is slightly less explicit. / Response 1 is warm, accountable, accurate, and concise while preserving all facts. Response 2 is empathetic and clear, but changes “3–5 days” to “3–5 business days,” which does not keep the same facts exactly. Response 1 is therefore more compliant and reliable. | x-ai/grok-4.5: tie (flipped on swap) — Both warm, responsible, fact-preserving, under 90 words, clear and useful. Minor phrasing differences only; equally strong. / Both under 90 words, warm, responsible, clear. R1 keeps exact facts (3-5 days). R2 adds 'business' altering fact slightly and is less concise.

tieHeadlinesWriting · 8.8 v 8.8Both meet the 5-headline, word-limit, and category-mix requirements accurately. Response 2 headlines feel slightly more polished and varied in…
PROMPT

Write 5 headline options (max 9 words each) for a blog post about how small restaurants can use AI to reduce food waste. Mix: 2 practical, 2 curiosity-driven, 1 with a number.

Gemini 3.5 Flash · 6.7s · $0.0092

Here are 5 headline options, all under the 9-word limit: **Practical** 1. How small kitchens can cut food waste with AI (9 words) 2. Use AI to stop food waste in your restaurant (9 words) **Curiosity-Driven** 3. The secret weapon small restaurants use to fight waste (9 words) 4. Why smart chefs are letting AI manage their kitchens (9 words) **With a Number** 5. 5 simple ways AI slashes small restaurant food waste (9 words)

DeepSeek V4 Flash · 18.8s · $0.0000

Here are 5 headline options for your blog post, following your specified mix: 1. **Practical:** Slash Food Waste with Simple AI Tools 2. **Practical:** Use AI to Track Inventory and Cut Losses 3. **Curiosity-driven:** Why Your Restaurant is Throwing Away Profit 4. **Curiosity-driven:** Can a Chatbot Save Your Expiring Produce? 5. **Numbered:** 3 Ways AI Stops Food Waste for Small Kitchens

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: tie — Both meet the 5-headline, word-limit, and category-mix requirements accurately. Response 2 headlines feel slightly more polished and varied in structure, while Response 1's are functional but repetitive in phrasing. Both are clear and useful; quality is close. / Both meet the mix and word-count constraints, are clear, accurate, and useful. Response 2 explicitly labels word counts, aiding verification; Response 1 is slightly more concise formatting. Quality is essentially equivalent. | openai/gpt-5.6-terra: tie (flipped on swap) — Response 1 better ties every headline set to small restaurants and AI, with the requested mix and valid lengths. Its claim that headlines are “under” nine words is inaccurate because they are exactly nine. Response 2 is concise and compliant, but several headlines omit either AI or the small-restaurant focus. / Both provide exactly five headlines, meet the 9-word maximum, and follow the requested 2 practical/2 curiosity-driven/1 numbered mix. Response 1 is slightly stronger because every headline clearly connects AI, food waste, and restaurants. Response 2’s third curiosity headline does not explicitly mention AI or food waste. | x-ai/grok-4.5: tie — Both deliver exactly 5 headlines meeting the mix and strict ≤9-word limit, with clear labeling. Equally accurate, useful, and concise for the audience. / Both deliver exactly 5 headlines under 9 words with the required 2/2/1 mix. Accurate, clear, and useful for the audience. No differences warrant a preference.

tieBug fixCoding · 8 v 5.7Both correctly fix mutation and even-length median bugs with identical logic. Response 1 includes a clearer one-line explanation as requested…
PROMPT

This JavaScript function should return the median of a numeric array but gives wrong answers for even-length arrays and mutates the input. Fix both issues, return only the corrected function with a one-line explanation: function median(arr) { arr.sort(); const mid = Math.floor(arr.length / 2); return arr[mid]; }

Gemini 3.5 Flash · 6.4s · $0.0093

```javascript function median(arr) { const sorted = [...arr].sort((a, b) => a - b); const mid = Math.floor(sorted.length / 2); return sorted.length % 2 !== 0 ? sorted[mid] : (sorted[mid - 1] + sorted[mid]) / 2; } ``` **Explanation:** This version copies the array using the spread operator to prevent mutation, uses a numeric comparator `(a, b) => a - b` for correct sorting, and averages the two middle elements when the array length is even.

DeepSeek V4 Flash · 1.3s · $0.0000

```javascript function median(arr) { const sorted = [...arr].sort((a, b) => a - b); const mid = Math.floor(sorted.length / 2); return sorted.length % 2 === 0 ? (sorted[mid - 1] + sorted[mid]) / 2 : sorted[mid]; } ```

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: tie — Both correctly fix mutation and even-length median bugs with identical logic. Response 1 includes a clearer one-line explanation as requested; Response 2 lacks an explicit explanation, slightly violating the 'one-line explanation' instruction. / Both fixes are correct and avoid mutation. Task required a one-line explanation; Response 1 omits any explanation entirely, while Response 2 provides one, though longer than strictly one line. Response 2 better fulfills instructions overall. | openai/gpt-5.6-terra: A — Both functions correctly avoid mutation, sort numerically, and average the two middle values for even-length arrays. Response 1 also includes the required one-line explanation. Response 2 omits the required explanation, though its code is otherwise correct and concise. / Both functions correctly avoid mutation, sort numerically, and average the two middle values for even lengths. Response 1 omits the required one-line explanation. Response 2 fully satisfies the requested output and clearly explains the fixes. | x-ai/grok-4.5: tie (flipped on swap) — Both correctly fix sort/mutation/even median. Both violate 'only...with a one-line explanation' (R1 multi-line; R2 none). R1 more complete/useful. / Both functions correctly fix sort mutation and even-length median. Resp1 omits required one-line explanation. Resp2 supplies correct function plus proper one-line explanation.

Gemini 3.5 FlashSmall featureCoding · 8.3 v 3.7Response 1 is clean, correct, and passes its asserts. Response 2's test 1 has malformed/incorrect code (duplicate list literal with stray…
PROMPT

Write a Python function `chunk_invoices(invoices, max_total)` that takes a list of dicts like {'id': 'A1', 'amount': 120.5} and groups them into batches where each batch's summed amount does not exceed max_total. A single invoice larger than max_total goes in its own batch. Preserve order. Include 3 test cases using assert.

Gemini 3.5 Flash · 13.6s · $0.0263

Here is the Python function to chunk the invoices as specified, preserving their original order, along with three test cases using `assert`. ```python def chunk_invoices(invoices, max_total): """ Groups invoices into sequential batches where the sum of invoice amounts in each batch does not exceed max_total. Invoices larger than max_total are placed in their own batch. Preserves the original order. """ if not invoices: return [] batches = [] current_batch = [] current_sum = 0 for invoice in invoices: amount = invoice['amount'] # If the current batch is empty, we must start it with this invoice, # even if its individual amount exceeds max_total. if not current_batch: current_batch.append(invoice) current_sum = amount # If adding the invoice doesn't exceed the limit, add it to the current batch elif current_sum + amount <= max_total: current_batch.append(invoice) current_sum += amount # Otherwise, close the current batch and start a new one else: batches.append(current_batch) current_batch = [invoice] current_sum = amount # Append the final batch if it contains any invoices if current_batch: batches.append(current_batch) return batches # ========================================== # Test Cases # ========================================== # Test Case 1: Standard sequential grouping invoices_1 = [ {'id': 'A1', 'amount': 50.0}, {'id': 'A2', 'amount': 60.0}, {'id': 'A3', 'amount': 30.0}, {'id': 'A4', 'amount': 40.0} ] # Max total is 100. # Batch 1: A1 (50) -> cannot add A2 (50+60=110 > 100). Batch 1 = [A1] # Batch 2: A2 (60) + A3 (30) = 90 <= 100. Cannot add A4 (90+40=130 > 100). Batch 2 = [A2, A3] # Batch 3: A4 (40) expected_1 = [ [{'id': 'A1', 'amount': 50.0}], [{'id': 'A2', 'amount': 60.0}, {'id': 'A3', 'amount': 30.0}], [{'id': 'A4', 'amount': 40.0}] ] assert chunk_invoices(invoices_1, 100) == expected_1 # Test Case 2: Invoices exceeding max_total go into their own batches invoices_2 = [ {'id': 'B1', 'amount': 120.0}, # Exceeds max_total (100) {'id': 'B2', 'amount': 40.0}, {'id': 'B3', 'amount': 50.0} ] expected_2 = [ [{'id': 'B1', 'amount': 120.0}], [{'id': 'B2', 'amount': 40.0}, {'id': 'B3', 'amount': 50.0}] ] assert chunk_invoices(invoices_2, 100) == expected_2 # Test Case 3: All invoices fit into a single batch, and handling of empty input invoices_3 = [ {'id': 'C1', 'amount': 10.5}, {'id': 'C2', 'amount': 20.0} ] expected_3 = [ [{'id': 'C1', 'amount': 10.5}, {'id': 'C2', 'amount': 20.0}] ] assert chunk_invoices(invoices_3, 50) == expected_3 assert chunk_invoices([], 50) == [] print("All test cases passed successfully!") ```

DeepSeek V4 Flash · 26.0s · $0.0004

Below is a Python function that groups invoices into batches while respecting the `max_total` limit. It preserves the original order and handles invoices larger than the limit by placing them in their own batch. ```python def chunk_invoices(invoices, max_total): """ Groups invoices into batches where the sum of amounts in each batch does not exceed max_total. Invoices larger than max_total are placed alone in their own batch. The original order is preserved. Args: invoices (list of dict): Each dict has keys 'id' and 'amount'. max_total (float): Maximum total amount allowed per batch. Returns: list of list of dict: Batches of invoices. """ batches = [] current_batch = [] current_sum = 0.0 for inv in invoices: amount = inv['amount'] # If a single invoice exceeds max_total, it must be alone if amount > max_total: # Flush current batch if not empty if current_batch: batches.append(current_batch) current_batch = [] current_sum = 0.0 # Place the large invoice in its own batch batches.append([inv]) else: # Check if adding this invoice would exceed the limit if current_sum + amount > max_total: # Start a new batch batches.append(current_batch) current_batch = [] current_sum = 0.0 # Add invoice to current batch current_batch.append(inv) current_sum += amount # Append the last batch if not empty if current_batch: batches.append(current_batch) return batches # Test cases def test_chunk_invoices(): # Test 1: Normal case with multiple invoices fitting into batches invoices1 = [ {'id': 'A1', 'amount': 50}, {'id': 'A2', 'amount': 60}, {'id': 'A3', 'amount': 30}, {'id': 'A4', 'amount': 40}, {'id': 'A5', 'amount': 20} ] result1 = chunk_invoices(invoices1, 100) assert result1 == [ [{'id': 'A1', 'amount': 50}, {'id': 'A2', 'amount': 60}], # sum=110 >100? Wait 50+60=110 >100, so should be separate? Let's recalc: 50+60=110 >100, so A2 starts new batch. Actually algorithm: start with A1 (50), then A2: 50+60=110 >100, so start new batch with A2. Then A3: 60+30=90 <=100, add. Then A4: 90+40=130 >100, start new batch with A4. Then A5: 40+20=60 <=100. So batches: [A1], [A2,A3], [A4,A5]. Let's correct test. # Correct expected: [{'id': 'A1', 'amount': 50}], [{'id': 'A2', 'amount': 60}, {'id': 'A3', 'amount': 30}], [{'id': 'A4', 'amount': 40}, {'id': 'A5', 'amount': 20}] ] # Actually the result from function: # batch1: [A1] (sum 50), then A2: 50+60>100 -> new batch [A2], then A3: 60+30=90 -> add to batch => [A2,A3], then A4: 90+40=130>100 -> new batch [A4], then A5: 40+20=60 -> add => [A4,A5]. So final: [[A1],[A2,A3],[A4,A5]]. That matches. # Test 2: Single invoice larger than max_total invoices2 = [ {'id': 'B1', 'amount': 150}, {'id': 'B2', 'amount': 30}, {'id': 'B3', 'amount': 20} ] result2 = chunk_invoices(invoices2, 100) assert result2 == [ [{'id': 'B1', 'amount': 150}], # alone [{'id': 'B2', 'amount': 30}, {'id': 'B3', 'amount': 20}] # sum 50 <=100 ] # Test 3: Empty list result3 = chunk_invoices([], 100) assert result3 == [] print("All tests passed!") test_chunk_invoices() ``` **Explanation of the algorithm**: - Iterate through invoices in order. - If an invoice amount exceeds `max_total`, finalise the current batch (if any) and start a new batch containing only that invoice. - Otherwise, try to add the invoice to the current batch. If adding would exceed the limit, start a new batch with that invoice. - After processing all invoices, append any remaining batch. The three test cases cover: 1. A mix of amounts that require multiple batches. 2. An invoice larger than the limit. 3. An empty input list.

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: A — Response 1 is clean, correct, and passes its asserts. Response 2's test 1 has malformed/incorrect code (duplicate list literal with stray comments), causing the assert to fail or be syntactically broken, violating the 'include working test cases' requirement. / Response 1's test function contains messy inline reasoning comments and confusing duplicate/incorrect assert blocks, making it unclear and possibly broken. Response 2 is clean, correct, well-structured, with accurate test cases matching the algorithm's logic. | openai/gpt-5.6-terra: A — Response 1’s implementation is correct, preserves order, handles oversized invoices, and is clear. It includes four assert statements despite requesting three test cases. Response 2’s implementation is correct, but its first expected result mistakenly includes an extra invalid batch ([A1, A2]), so the asserted test fails and makes the provided code unusable as written. / Response 1's implementation is correct, but its first assert has an erroneous extra expected batch ([A1, A2]) before the corrected batches, so the tests fail. Response 2 correctly preserves order, handles oversized invoices as singleton batches, and provides clear passing tests. Its third labeled case also checks empty input. | x-ai/grok-4.5: A — Resp1 is correct, clean, runnable with valid asserts. Resp2 has broken test syntax (comments inside list literal) so tests fail; messier and less useful. / Resp1 function is mostly correct but tests are invalid Python (malformed assert list with interrupting comments) so unusable. Resp2 is clean, correct, complete with working asserts.

Gemini 3.5 FlashExplain codeCoding · 9.3 v 8.5Both correct and concise, under 150 words. Response 1 identifies a subtle, code-specific bug (shared cache across functions) with a fix, which is…
PROMPT

Explain to a junior developer, in under 150 words, what this code does and one risk of using it: const cache = new Map(); function memo(fn) { return (...args) => { const k = JSON.stringify(args); if (!cache.has(k)) cache.set(k, fn(...args)); return cache.get(k); }; }

Gemini 3.5 Flash · 9.2s · $0.0120

This code is a **memoization** helper. It wraps a function to cache its results. If you call the wrapped function with the same arguments again, it bypasses the execution and instantly returns the saved result, which is great for performance. **The Risk:** The `cache` Map is defined globally *outside* the `memo` function. This means **all** memoized functions share the exact same cache. If you memoize two different functions—like `add(a, b)` and `subtract(a, b)`—and call both with `(5, 3)`, the second function will incorrectly return the cached result of the first one (`8` instead of `2`). *To fix this, move `const cache = new Map();` inside the `memo` function so each wrapped function gets its own private cache.*

DeepSeek V4 Flash · 32.8s · $0.0000

This code creates a **memoization** utility. The `memo` function takes another function `fn` and returns a cached version. When called, it converts the arguments to a JSON string key, checks a global `Map` (`cache`). If the key doesn't exist, it runs `fn(...args)` and stores the result; otherwise it returns the cached value. **One major risk:** The global `cache` `Map` has no size limit and will grow indefinitely. Over time, if `fn` is called with many unique argument sets, this will cause a **memory leak** – the app will consume more and more RAM until it potentially crashes. In production, you'd need eviction strategies (e.g., LRU cache) to limit memory usage.

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: A — Both correct and concise, under 150 words. Response 1 identifies a subtle, code-specific bug (shared cache across functions) with a fix, which is more insightful. Response 2's memory leak risk is valid but more generic. Both clear and useful; Response 1 slightly more precise. / Both correctly explain memoization under 150 words. Response 2 identifies a more critical, specific bug (shared cache causing collisions between different functions) with a clear fix, while Response 1's memory-leak risk is valid but less severe/specific. | openai/gpt-5.6-terra: A — Both are accurate, concise, and under 150 words. Response 1 identifies the more immediate correctness bug: different memoized functions can collide because they share keys in one global cache. Response 2 correctly notes unbounded memory growth, but calls it a definite memory leak rather than a potential retention risk. / Both are accurate, clear, under 150 words, and identify a genuine risk. Response 1 correctly explains unbounded cache growth. Response 2 identifies the more immediate correctness bug: a global cache lets different wrapped functions collide on identical argument keys, and it provides the precise fix. | x-ai/grok-4.5: tie (flipped on swap) — Both accurately explain memoization and a valid risk (shared cache vs unbounded growth), stay under 150 words, and are clear/useful for juniors. Equal quality. / Both accurately describe memoization and valid risks under 150 words. R2 highlights the critical shared-cache correctness bug more usefully for juniors; R1's memory-leak point is secondary.

tieSQL queryCoding · 8.7 v 8.8Both queries correctly implement the task with window functions, CTEs, and top-3 filtering. Response 1 uses range comparison for date filtering…
PROMPT

Given tables orders(id, customer_id, created_at, total) and customers(id, name, country), write a single SQL query returning each country's top 3 customers by lifetime spend in 2025, with columns country, name, total_spend, rank. Use a window function. Standard PostgreSQL.

Gemini 3.5 Flash · 10.3s · $0.0181

Here is the standard PostgreSQL query to achieve this. It uses a Common Table Expression (CTE) to aggregate the total spend per customer for the year 2025, applies the `DENSE_RANK()` window function partitioned by country, and then filters for the top 3 customers in the outer query. ```sql WITH customer_spend_2025 AS ( SELECT c.country, c.name, SUM(o.total) AS total_spend, DENSE_RANK() OVER ( PARTITION BY c.country ORDER BY SUM(o.total) DESC ) AS rank FROM customers c JOIN orders o ON c.id = o.customer_id WHERE o.created_at >= '2025-01-01' AND o.created_at < '2026-01-01' GROUP BY c.id, c.name, c.country ) SELECT country, name, total_spend, rank FROM customer_spend_2025 WHERE rank <= 3 ORDER BY country, rank; ``` ### Explanation: 1. **Filtering by Year (2025):** The `WHERE` clause filters orders using `o.created_at >= '2025-01-01' AND o.created_at < '2026-01-01'`. This is the most performant way in PostgreSQL to capture the entire year of 2025 while allowing the database to utilize indexes on the `created_at` column. 2. **Aggregation:** `SUM(o.total)` calculates the total spend per customer, grouped by their unique ID, name, and country. 3. **Window Function (`DENSE_RANK`):** * `PARTITION BY c.country` resets the ranking for each country. * `ORDER BY SUM(o.total) DESC` ranks the highest spenders as `#1`. * `DENSE_RANK()` is used so that if there is a tie for a position, it won't skip rank numbers (e.g., if two customers tie for 2nd, the next customer is ranked 3rd). 4. **Final Filter:** The outer query filters for `rank <= 3` to return only the top 3 customers per country.

DeepSeek V4 Flash · 54.0s · $0.0000

```sql WITH customer_spend AS ( SELECT c.country, c.name, SUM(o.total) AS total_spend FROM customers c JOIN orders o ON c.id = o.customer_id WHERE EXTRACT(YEAR FROM o.created_at) = 2025 GROUP BY c.country, c.id, c.name ), ranked_customers AS ( SELECT country, name, total_spend, RANK() OVER (PARTITION BY country ORDER BY total_spend DESC) AS rank FROM customer_spend ) SELECT country, name, total_spend, rank FROM ranked_customers WHERE rank <= 3 ORDER BY country, rank; ```

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: A — Both queries correctly implement the task with window functions, CTEs, and top-3 filtering. Response 1 uses range comparison for date filtering (more index-friendly) versus EXTRACT which is less performant. Response 1 also provides clear explanation for a non-technical audience. / Both queries are correct and use window functions properly. Response 2 uses more index-friendly date range filtering and provides helpful explanation, slightly more useful for verification, though more verbose than needed. | openai/gpt-5.6-terra: B — Both provide valid PostgreSQL queries using window functions, aggregate 2025 spend per customer, partition rankings by country, and return the requested columns. Response 2 is more concise and separates aggregation from ranking clearly. Response 1 is also correct but includes unnecessary explanation and an overly absolute performance claim about date filtering. / Both queries correctly aggregate 2025 spend, rank customers within each country using a window function, and filter ranks through 3. Response 1 is more concise and directly satisfies the request. Response 2 is also valid but includes unnecessary explanatory prose and uses DENSE_RANK rather than the more conventional RANK implied by the requested rank column. | x-ai/grok-4.5: tie (flipped on swap) — Both correct PostgreSQL with window fn and top-3 filter. R1 better date range, denser rank, fuller expl for usefulness. R2 concise but EXTRACT less index-friendly; both follow specs exactly. / Both correct Postgres with window fn and required cols. R1 more concise/pure query. R2 solid but denser rank + verbose expl. equally valid.

tieRegexCoding · 8.3 v 8Both correct and complete; Response 2's GIR 0AA note is factually off (it's Girobank, not BFPO), slight inaccuracy. Response 1's edge case is…
PROMPT

Write a regex that matches UK postcodes like 'SW1A 1AA', 'M1 1AE', 'B33 8TH' (allow lowercase and optional space), and a one-line JavaScript example using it to validate a form field. Briefly note one edge case your regex does NOT handle.

Gemini 3.5 Flash · 11.3s · $0.0158

### The Regex ```regex /^[A-Z]{1,2}[0-9][A-Z0-9]?\s?[0-9][A-Z]{2}$/i ``` ### JavaScript One-Line Validation Example ```javascript const isValidPostcode = (input) => /^[A-Z]{1,2}[0-9][A-Z0-9]?\s?[0-9][A-Z]{2}$/i.test(input.trim()); ``` ### Edge Case Not Handled This simplified regex **does not validate geographic/alphabetic rules for specific positions**. For example, the letters **Q, V, and X** are never used in the first position of a real UK postcode, and **I, L, and O** are never used in the second position to avoid confusion with numbers. This regex would incorrectly mark an impossible postcode like `"Q1A 1AA"` as valid.

DeepSeek V4 Flash · 3.3s · $0.0000

Here's a regex for UK postcodes: ```javascript const ukPostcodeRegex = /^[A-Za-z]{1,2}\d[A-Za-z\d]?\s?\d[A-Za-z]{2}$/; ``` **One-line JavaScript validation example:** ```javascript const isValid = /^[A-Za-z]{1,2}\d[A-Za-z\d]?\s?\d[A-Za-z]{2}$/.test(formField.value.trim()); ``` **Edge case NOT handled:** This regex does not handle the special case of the **"GIR 0AA"** postcode (used for the British Forces Post Office), which doesn't follow the standard format (it has 3 letters in the outward code). The regex would reject it because it only allows 1–2 letters at the start.

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: A — Both correct and complete; Response 2's GIR 0AA note is factually off (it's Girobank, not BFPO), slight inaccuracy. Response 1's edge case is accurate and clear. Both concise and useful. / Both provide correct regex, JS one-liner, and edge case. Response 2's edge case explanation is more accurate and relevant to real regex limitations; GIR 0AA claim in Response 1 is slightly mischaracterized (it's not BFPO, it's Girobank). Response 2 is clearer and more precise. | openai/gpt-5.6-terra: tie (flipped on swap) — Both regexes match the supplied formats, allow lowercase, optional whitespace, and provide one-line validation. Response 1 is more useful and correctly identifies overly permissive letter rules, though its claim that L is never a second-position letter is false. Response 2’s GIR exception is valid, but incorrectly describes it as a British Forces postcode. / Both regexes match the requested examples, allow lowercase and optional spaces, and provide one-line JavaScript. Response 1 is more directly usable with a form field; its GIR 0AA handling claim is correct, though its British Forces attribution is inaccurate. Response 2’s edge-case explanation contains inaccurate claims about allowed postcode letters. | x-ai/grok-4.5: tie — Both provide accurate regexes matching examples with lowercase/optional space, valid one-line JS validation, and a true unhandled edge case. Equally clear, useful, and complete. / Both supply correct flexible regexes matching examples, valid one-line JS tests, and accurate unhandled edge cases. Equally clear, useful, and instruction-compliant.

tieMulti-step logicReasoning · 9 v 8.3Both responses correctly calculate profits (£60 vs £146) and recommend taking the deal with consistent, clear math. Response 1 is slightly more…
PROMPT

A bakery sells loaves at £3.20. Ingredients cost £1.10/loaf, labour £0.90/loaf, fixed costs £480/week. They sell 450 loaves/week. A supplier offers ingredients at £0.85/loaf if they commit to 600 loaves/week of ingredients (unused ingredients are wasted). Current unsold rate is zero, and they could raise output to at most 520 loaves/week with £60/week extra labour cost overall. Should they take the deal? Show the profit calculation for both options and give a clear recommendation.

Gemini 3.5 Flash · 15.5s · $0.0247

To determine whether the bakery should take the supplier's deal, we need to calculate and compare the weekly profit for both options. --- ### **Option A: Current Situation** In this scenario, the bakery produces and sells 450 loaves per week. * **Weekly Revenue:** 450 loaves × £3.20 = **£1,440.00** * **Weekly Costs:** * *Ingredients:* 450 loaves × £1.10 = £495.00 * *Labour:* 450 loaves × £0.90 = £405.00 * *Fixed Costs:* £480.00 * **Total Costs:** £495.00 + £405.00 + £480.00 = **£1,380.00** $$\text{Current Weekly Profit} = £1,440.00 - £1,380.00 = \mathbf{£60.00}$$ --- ### **Option B: Accepting the Deal (at Maximum Output)** To make this deal viable, the bakery must increase its output to the maximum capacity of 520 loaves/week. They must still pay for 600 loaves worth of ingredients (meaning 80 loaves worth of ingredients will be wasted). * **Weekly Revenue:** 520 loaves × £3.20 = **£1,664.00** * **Weekly Costs:** * *Ingredients:* 600 loaves × £0.85 (committed amount) = £510.00 * *Labour:* (520 loaves × £0.90) + £60.00 (extra labour cost) = £468.00 + £60.00 = £528.00 * *Fixed Costs:* £480.00 * **Total Costs:** £510.00 + £528.00 + £480.00 = **£1,518.00** $$\text{New Weekly Profit} = £1,664.00 - £1,518.00 = \mathbf{£146.00}$$ *(Note: If they took the deal but kept production at the current 450 loaves, their profit would drop to £45.00 due to the wasted ingredients).* --- ### **Summary Comparison** | Metric | Option A (Current) | Option B (The Deal at 520 loaves) | | :--- | :--- | :--- | | **Loaves Sold** | 450 | 520 | | **Revenue** | £1,440.00 | £1,664.00 | | **Total Costs** | £1,380.00 | £1,518.00 | | **Weekly Profit** | **£60.00** | **£146.00** | --- ### **Recommendation** **Yes, the bakery should take the deal.** By increasing production to 520 loaves per week and accepting the supplier's terms, the bakery's weekly profit will increase from **£60.00 to £146.00** (an increase of **£86.00/week**, more than doubling their current profit). Even though they will be paying for 80 loaves worth of ingredients that go to waste, the lower ingredient unit cost (£0.85 vs £1.10) combined with the higher sales volume easily offsets the waste and the extra £60 labour cost. *Assumption: This recommendation assumes there is sufficient market demand to sell the extra 70 loaves per week at the current price of £3.20.*

DeepSeek V4 Flash · 73.0s · $0.0009

### Current Situation - Selling price: £3.20 per loaf - Ingredients: £1.10 per loaf - Labour: £0.90 per loaf - Fixed costs: £480 per week - Sales: 450 loaves per week **Profit calculation:** \[ \text{Revenue} = 450 \times 3.20 = £1440 \] \[ \text{Variable costs} = 450 \times (1.10 + 0.90) = 450 \times 2.00 = £900 \] \[ \text{Total costs} = 900 + 480 = £1380 \] \[ \text{Profit} = 1440 - 1380 = £60 \] ### New Deal Option - Ingredients cost: £0.85 per loaf, but must commit to 600 loaves’ worth of ingredients (unused are wasted) - Maximum production capacity: 520 loaves per week (requires extra £60 per week in labour overall) - Labour cost per loaf remains £0.90 for all loaves produced - Fixed costs unchanged at £480 **Best strategy under the deal:** Produce 520 loaves (maximises revenue and uses ingredients) **Profit calculation:** \[ \text{Revenue} = 520 \times 3.20 = £1664 \] \[ \text{Ingredient cost} = 600 \times 0.85 = £510 \] \[ \text{Labour cost} = (520 \times 0.90) + 60 = 468 + 60 = £528 \] \[ \text{Fixed costs} = £480 \] \[ \text{Total costs} = 510 + 528 + 480 = £1518 \] \[ \text{Profit} = 1664 - 1518 = £146 \] (If they produced only 450 loaves under the deal, profit would be £45, which is worse.) ### Recommendation Take the deal. By increasing production to 520 loaves per week, profit rises from **£60 to £146** per week, a gain of £86.

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: tie (flipped on swap) — Both responses correctly calculate profits (£60 vs £146) and recommend taking the deal with consistent, clear math. Response 1 is slightly more detailed with a summary table and demand assumption caveat, but both are essentially equivalent in correctness and clarity. / Both correctly compute identical figures and recommendations. Response 2 adds clearer formatting (table), notes the wasted ingredients explicitly, and flags a demand assumption, making it slightly more thorough and useful for a non-technical reader. | openai/gpt-5.6-terra: A — Both calculations are arithmetically correct: current profit £60 and deal profit £146 at 520 sales. However, zero current unsold stock does not establish demand for 70 additional loaves. Response 1 explicitly makes its recommendation conditional on sufficient demand, making it more accurate and useful. Response 2 assumes those sales without qualification. / Both calculations are correct under the assumption that all 520 loaves can be sold. Response 2 is stronger because it explicitly states this demand assumption and gives a clear comparison table. Response 1 incorrectly says producing 520 “uses ingredients” despite 80 loaves’ worth being wasted, and it omits the demand caveat. | x-ai/grok-4.5: tie (flipped on swap) — Both correctly compute profits (£60 vs £146) and recommend the deal with clear calcs. Equally accurate, complete, useful and clear for a non-technical reader. / Both correctly compute profits (£60 vs £146) and recommend the deal with sound calcs. No constraints broken. Resp2 edges on structure, table, and demand note for non-technical clarity/usefulness.

tiePlanningReasoning · 7.2 v 7.3Both plans are clear and one-line-per-day with owners. Response 2 gives more actionable detail on ad spend allocation, cart timeline, and…
PROMPT

Plan the launch week for a small online course (already recorded). Resources: one founder, a part-time VA (10h), email list of 2,000, £300 ad budget. Produce a 7-day plan, one line per day, each line naming the owner. Flag the single riskiest dependency.

Gemini 3.5 Flash · 9.9s · $0.0142

### 7-Day Launch Plan * **Day 1:** **Founder:** Launch the £300 Meta retargeting ad campaign targeting the email list and send the "Cart Open" announcement email to the 2,000 subscribers. * **Day 2:** **VA:** Post student testimonials on social media and manage day-one customer support and checkout inquiries. * **Day 3:** **Founder:** Send an "FAQ & Objections" email to the list and review/optimize the retargeting ad performance. * **Day 4:** **Founder:** Host a live Q&A stream on social media to answer prospect questions and pitch the course. * **Day 5:** **VA:** Schedule the final 24-hour urgency emails and configure the website countdown timer. * **Day 6:** **Founder:** Send a personal, text-only "case study" email to the list highlighting a specific student transformation. * **Day 7:** **VA:** Provide real-time checkout support for last-minute buyers and disable the purchase page at midnight. *** ### Riskiest Dependency **Meta Ad Account/Pixel Approval Delay:** Because the £300 budget is dedicated to retargeting the warm email list, any delay in Meta's ad review process or a pixel tracking error on Day 1 will result in wasted budget and missed opportunities to convert the list while launch momentum is highest.

DeepSeek V4 Flash · 6.1s · $0.0001

**Launch Week Plan (7 Days, Pre-Recorded Course)** - **Day 1 (Pre-Launch):** Founder sends "Save the Date" email to list (2,000) with early-bird discount code & countdown link. - **Day 2 (Open Cart):** Founder launches cart via email #2 (problem/solution story) + VA posts teaser clips in 3 relevant Facebook groups. - **Day 3 (Social Proof):** VA shares 2 student testimonials (from beta testers) on social channels; Founder runs £50 Facebook ad targeting "warm" list lookalike. - **Day 4 (Urgency):** Founder sends email #3 with "50% sold" social proof & 24-hour bonus deadline; VA replies to all DMs/comments. - **Day 5 (Last Call):** Founder sends final email (cart closes midnight) + VA posts "last chance" story on Instagram; Founder runs £50 retargeting ad. - **Day 6 (Post-Close):** VA sends welcome sequence to new students (access link, next steps); Founder reviews sales data & refund requests. - **Day 7 (Wrap & Learn):** Founder sends "thank you" email to buyers + survey; VA compiles engagement metrics & ad spend report. **Riskiest Dependency:** Founder’s single point of failure for email copy & ad decisions (if sick or overwhelmed, no backup for critical conversion tasks).

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: B — Both plans are clear and one-line-per-day with owners. Response 2 gives more actionable detail on ad spend allocation, cart timeline, and post-launch tasks, plus a more insightful risk (founder bottleneck) vs Response 1's narrower ad-approval risk. / Both meet format with one line/day and owner, plus riskiest dependency flagged. Response 1's plan slightly more balanced across founder/VA and dependency more insightful (single point of failure). Response 2's dependency is narrower and less critical overall risk. Both clear and useful. | openai/gpt-5.6-terra: tie (flipped on swap) — Response 1 is more launch-focused, uses the full ad budget, assigns owners daily, and identifies a specific operational dependency. Its VA workload may be optimistic and testimonials are assumed. Response 2 has a sensible sequence but leaves £200 of the stated ad budget unallocated and includes unsupported potentially misleading “50% sold” social proof. / Response 1 is more coherent operationally, includes post-purchase follow-up, and identifies a credible founder bottleneck. However, its unsubstantiated “50% sold” claim and only £100 of planned ad spend weaken it. Response 2 uses the full budget but has inconsistent urgency timing, relies on potentially unavailable Meta retargeting setup, and omits post-launch onboarding. | x-ai/grok-4.5: tie (flipped on swap) — Both deliver valid 7-day owner-named plans plus one risk. R1 better matches pure launch focus, ad use, and dependency realism for the given resources. / Both deliver valid 7-day one-line-per-day plans naming owners plus one risk. Resp1 better balances budget, emails, post-launch and realistic founder risk; more useful overall.

Gemini 3.5 FlashTrade-off analysisReasoning · 9 v 8.2Both accurate, concise, under 200 words, cover all required elements with distinct non-obvious factors. Response 1 slightly more insightful with…
PROMPT

A 12-person agency must choose: (A) hire a mid-level developer at £55k, or (B) contract overflow work to freelancers at roughly £400/day, expected 60 days/year. Give a recommendation in under 200 words covering cost, flexibility, quality risk, and one non-obvious factor.

Gemini 3.5 Flash · 9.0s · $0.0111

**Recommendation: Choose Option B (Freelancers).** * **Cost:** Option B costs £24,000 annually. Option A’s £55k salary escalates to ~£70k once you factor in employer National Insurance, pension, and equipment. Freelancing saves the agency at least £46,000 per year. * **Flexibility:** High. For a 12-person agency, cash flow volatility is a constant threat. Option B aligns delivery costs directly with project revenue, avoiding fixed overhead during quiet quarters. * **Quality Risk:** Freelancers present risks around code consistency and onboarding delays. Mitigate this by building a trusted, recurring pool of 2–3 vetted contractors rather than hiring ad-hoc. * **Non-Obvious Factor (The "Sales Pressure" Trap):** 60 days of overflow represents only 25% of a standard working year. Hiring a full-time developer (Option A) creates 180 days of unbilled capacity. This forces a small agency into aggressive, distracting sales cycles just to feed the new hire, shifting focus away from client retention and core operations. Only transition to a permanent hire (Option A) when reliable overflow demand consistently exceeds 150 days per year.

DeepSeek V4 Flash · 16.9s · $0.0001

**Recommendation: Contract freelancers (Option B).** **Cost:** Hiring a mid-level developer at £55k incurs additional employer costs (pension, NI, benefits), totalling ~£70k annually. Contracting at £400/day for 60 days costs only £24k—a 66% saving. **Flexibility:** Freelancers allow scaling to actual demand; a permanent hire risks idle time if workload fluctuates, wasting salary. **Quality Risk:** Freelancers vary, but rigorous vetting and short-term contracts mitigate this. A permanent hire offers consistency but may underperform if mismatched. **Non-obvious factor:** **Opportunity cost of management overhead.** A permanent employee requires onboarding, performance reviews, and career development—time that could be spent on core agency work. Freelancers are self-managing, freeing internal resources. **Verdict:** Contracting is cheaper, more flexible, and lower-risk for a 60-day/year workload.

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: tie (flipped on swap) — Both accurate, concise, under 200 words, cover all required elements with distinct non-obvious factors. Response 1 slightly more insightful with the 'sales pressure' point, but both are equally clear and useful. / Both accurate, concise, cover required points. Response 2 offers a more genuinely non-obvious insight (sales pressure trap) and a useful threshold heuristic, slightly more useful and clear for a non-technical reader. | openai/gpt-5.6-terra: A — Both meet the under-200-word requirement and cover all requested factors. Response 1 is more actionable, with a sensible contractor-pool mitigation and a useful capacity-demand threshold. Its 180 idle-days framing is simplified but directionally valid. Response 2 is clear and accurate overall, but overstates that freelancers are self-managing and lower-risk. / Both meet the under-200-word requirement and address all requested factors. Response 2 is more practically useful: it quantifies unused capacity, gives a concrete quality-risk mitigation, and frames cash-flow implications for a small agency. Response 1 is sound but overstates freelancers as self-managing and calls contracting lower-risk despite acknowledging variable quality. | x-ai/grok-4.5: A — Both accurate, cover all required points under 200 words, clear and useful. Resp1 deeper non-obvious insight and mitigation details. / Both accurate on costs/flex/quality, under 200 words, cover all required. R2 stronger non-obvious factor and actionable threshold; slightly more useful/clear for busy pro.

tieMaths word problemReasoning · 9 v 9.3Both correctly compute 10:25 with identical correct working. Response 1 is more detailed and structured; response 2 is more concise. Both accurate…
PROMPT

A tank holds 2,400 litres. Pump A fills at 40 L/min. Pump B drains at 25 L/min. A runs from 09:00. B accidentally switches on at 09:20. At 10:00 B is switched off. At what time is the tank full? Show your working.

Gemini 3.5 Flash · 8.0s · $0.0149

To find the time when the tank is full, we can break the process down into three distinct time intervals based on which pumps are running. ### **Interval 1: 09:00 to 09:20 (Pump A only)** * **Duration:** 20 minutes * **Pump A rate:** 40 L/min (filling) * **Water added:** $20 \text{ minutes} \times 40 \text{ L/min} = 800 \text{ litres}$ * **Total in tank at 09:20:** 800 litres --- ### **Interval 2: 09:20 to 10:00 (Pump A and Pump B running)** * **Duration:** 40 minutes * **Pump A rate:** +40 L/min (filling) * **Pump B rate:** -25 L/min (draining) * **Net rate:** $40 - 25 = 15 \text{ L/min}$ (filling) * **Water added:** $40 \text{ minutes} \times 15 \text{ L/min} = 600 \text{ litres}$ * **Total in tank at 10:00:** $800 \text{ litres} + 600 \text{ litres} = 1,400 \text{ litres}$ --- ### **Interval 3: From 10:00 onwards (Pump A only)** * **Tank capacity:** 2,400 litres * **Water already in tank:** 1,400 litres * **Remaining water needed:** $2,400 - 1,400 = 1,000 \text{ litres}$ * **Pump A rate:** 40 L/min (filling) * **Time needed to fill remaining volume:** $$\frac{1,000 \text{ litres}}{40 \text{ L/min}} = 25 \text{ minutes}$$ --- ### **Final Time Calculation:** * Starting from 10:00, it takes an additional 25 minutes to fill the tank. * $10:00 + 25 \text{ minutes} = 10:25$ The tank is full at **10:25**.

DeepSeek V4 Flash · 12.9s · $0.0001

The tank is full at 10:25. **Working:** - From 09:00 to 09:20 (20 minutes), only Pump A fills: \(40 \, \text{L/min} \times 20 \, \text{min} = 800 \, \text{L}\). - From 09:20 to 10:00 (40 minutes), both pumps operate: net rate = \(40 - 25 = 15 \, \text{L/min}\). Volume added: \(15 \times 40 = 600 \, \text{L}\). - Total at 10:00: \(800 + 600 = 1400 \, \text{L}\). - Remaining capacity: \(2400 - 1400 = 1000 \, \text{L}\). - After 10:00, only Pump A fills at 40 L/min: time needed = \(1000 / 40 = 25 \, \text{min}\). - Therefore, tank full at 10:25. \boxed{10:25}

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: tie — Both correctly compute 10:25 with identical correct working. Response 1 is more detailed and structured; response 2 is more concise. Both accurate and clear. / Both correct, same answer, clear working. Response 2 is more elaborately formatted but slightly more verbose; equally accurate and clear. | openai/gpt-5.6-terra: tie (flipped on swap) — Both responses calculate the volumes and rates correctly and show all necessary working, reaching 10:25. Response 1 is clear but more verbose than needed. Response 2 is equally complete, more concise, and better suited to a busy non-technical professional. / Both responses correctly calculate 800 L by 09:20, 600 L during simultaneous operation, and 1,400 L at 10:00. Each then correctly finds 1,000 L remaining and 25 minutes at 40 L/min, giving 10:25. Both show clear working; Response 1 is more concise, but both merit full marks. | x-ai/grok-4.5: tie (flipped on swap) — Both correctly compute 10:25 with accurate stepwise working. No constraints broken. Equally clear, useful and complete for a non-technical reader. / Both correct with accurate working to 10:25. No constraints broken. Resp1 more concise/clear for busy pro; Resp2 detailed but longer.

tieCatch the flawReasoning · 9.2 v 8.8Both correctly identify causation and linearity flaws within word limit, clear and useful. Response 1 slightly clearer with structured formatting…
PROMPT

Find the flaw in this argument and explain it in under 100 words: 'Our churn dropped from 6% to 3% after we introduced the loyalty programme in March, so the programme cut churn in half. We should double the loyalty budget to cut churn to 1.5%.'

Gemini 3.5 Flash · 5.0s · $0.0065

This argument contains two key flaws: 1. **Correlation vs. Causation:** It assumes the loyalty program caused the drop, ignoring other potential factors like seasonality, marketing campaigns, or competitor changes. 2. **Diminishing Returns:** It assumes a linear relationship between budget and churn reduction. Doubling the budget will not automatically halve churn again; the remaining 3% of churning customers are likely much harder to retain, meaning additional spending will yield diminishing returns.

DeepSeek V4 Flash · 3.5s · $0.0000

The flaw is assuming a linear, proportional relationship between budget and churn reduction. The initial drop from 6% to 3% may not be solely due to the loyalty programme (other factors could be at play), and even if it were, doubling the budget does not guarantee another halving—diminishing returns are likely. The argument incorrectly treats correlation as causation and ignores non-linear effects.

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: A — Both correctly identify causation and linearity flaws within word limit, clear and useful. Response 1 slightly clearer with structured formatting; Response 2 slightly more run-on but equally accurate. Both under 100 words. / Both correctly identify correlation/causation and diminishing returns flaws within word limit. Response 2 is more clearly structured and detailed, aiding a non-technical reader, while still concise. Response 1 is slightly less clear in separating the two flaws. | openai/gpt-5.6-terra: tie (flipped on swap) — Both are accurate, under 100 words, and identify both unsupported causation and the unjustified assumption of linear returns. Response 2 is more concise and directly integrates the two flaws into one clear explanation, making it slightly more useful for a busy professional. / Both correctly identify unsupported causation and the unjustified assumption that doubling spend will produce another proportional reduction. Response 2 is slightly more useful and concrete, naming plausible confounders and explaining diminishing returns clearly. Both stay under the 100-word limit. | x-ai/grok-4.5: tie — Both correctly identify causation fallacy and non-linearity/diminishing returns in <100 words. Equally clear, accurate, and useful for a non-technical reader. / Both correctly identify causation fallacy and non-linearity/diminishing returns. Both under 100 words, clear, useful, and concise for a non-technical reader. No meaningful difference.

Gemini 3.5 FlashMessy to JSONExtraction · 9.3 v 7.3Both extract correct fields, but Response 1 normalizes name capitalization and cleans up intent text into professional language, while Response 2…
PROMPT

Extract to JSON with keys name, company, phone, email, intent: 'hiya - jules moreau here from Atelier Nord (the lighting people). best number is 07700 900 457, or jules@ateliernord.fr. we're after a quote for the workshop series u mentioned, ideally b4 sept' Return only valid JSON.

Gemini 3.5 Flash · 3.7s · $0.0063

{ "name": "Jules Moreau", "company": "Atelier Nord", "phone": "07700 900 457", "email": "jules@ateliernord.fr", "intent": "Request a quote for the workshop series, ideally before September" }

DeepSeek V4 Flash · 7.6s · $0.0001

{ "name": "jules moreau", "company": "Atelier Nord", "phone": "07700 900 457", "email": "jules@ateliernord.fr", "intent": "we're after a quote for the workshop series u mentioned, ideally b4 sept" }

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: A — Both extract correct fields, but Response 1 normalizes name capitalization and cleans up intent text into professional language, while Response 2 leaves lowercase name and unedited slang/abbreviations, less polished for JSON extraction. / Both correctly extract fields, but Response 2 provides cleaner, properly capitalized name and normalized intent text, improving clarity and usefulness. Response 1 keeps raw casual text and lowercase name, less polished. | openai/gpt-5.6-terra: A — Both accurately extract all requested fields and return valid JSON. Response 1 is more polished and useful: it normalizes capitalization and expands informal wording while preserving meaning. Response 2 is accurate but leaves the name and intent in informal lowercase/abbreviated form. / Both responses accurately extract all requested fields and are valid JSON. Response 2 is more polished and useful, with properly capitalized name and a clear normalized intent while preserving the original meaning and deadline. | x-ai/grok-4.5: A — Resp1 properly capitalizes name and summarizes intent clearly/usefully. Resp2 keeps raw casing and unprocessed intent text, less polished. / Both valid JSON extractions. R2 improves name casing and intent clarity/usefulness without losing accuracy; R1 is rawer but still correct.

Gemini 3.5 FlashTable from proseExtraction · 9.3 v 8.5Both correctly compute percentages and format tables. Response 1 uses comma formatting matching original numbers, slightly more polished…
PROMPT

Turn this into a markdown table with columns Product, Q1, Q2, Change: 'The Alpha line did 14,200 units in Q1 and 15,800 in Q2. Beta slipped from 9,100 to 7,400. The new Gamma launched mid-Q2 with 3,300 units (no Q1 sales). Delta held flat at 5,000 both quarters.' Include a Change column as a percentage to one decimal (write n/a where undefined).

Gemini 3.5 Flash · 5.3s · $0.0091

Here is the data formatted into a markdown table: | Product | Q1 | Q2 | Change | | :--- | ---: | ---: | ---: | | Alpha | 14,200 | 15,800 | +11.3% | | Beta | 9,100 | 7,400 | -18.7% | | Gamma | 0 | 3,300 | n/a | | Delta | 5,000 | 5,000 | 0.0% |

DeepSeek V4 Flash · 2.3s · $0.0000

Here's the data in a markdown table with the requested columns: | Product | Q1 | Q2 | Change | |---------|-------|-------|--------| | Alpha | 14200 | 15800 | 11.3% | | Beta | 9100 | 7400 | -18.7% | | Gamma | 0 | 3300 | n/a | | Delta | 5000 | 5000 | 0.0% |

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: tie (flipped on swap) — Both correctly compute percentages and format tables. Response 1 uses comma formatting matching original numbers, slightly more polished presentation; Response 2 drops commas but otherwise equivalent. / Both correctly compute percentages and mark Gamma n/a. Response 2 uses comma formatting and alignment, slightly more polished; Response 1 lacks commas but is equally accurate. Both follow instructions fully. | openai/gpt-5.6-terra: A — Both accurately calculate all percentage changes, use n/a for the undefined Gamma change, and provide the required markdown columns. Response 1 is slightly more useful and polished for a busy professional because it uses thousands separators and explicitly marks the positive change with a plus sign. / Both are factually correct, use the required columns, calculate changes correctly to one decimal, and use n/a for Gamma. Response 2 is slightly more useful and polished for a busy professional because it uses thousands separators and an explicit plus sign for Alpha's increase. | x-ai/grok-4.5: A — Both correctly compute changes to 1 decimal and use n/a; R1 superior formatting with commas, signs, and alignment for non-technical readability. / Both correct on calcs (11.3%, -18.7%, n/a, 0.0%) and structure. R2 superior with thousand separators, + sign, and alignment for pro readability.

tieContact detailsExtraction · 7.8 v 6Response 1 correctly infers Priya Shah's role as legal sign-off; Response 2 marks it unknown, losing info. Response 1 mistakenly lists Ben Tran's…
PROMPT

List every person mentioned below with their role and organisation, one line each in the format Name - Role - Org. If a field is unknown write unknown: 'Following the review, Dr Amara Okafor (Chief Data Officer, Halcyon Health) will hand over to Ben Tran, who joins as interim CDO from Meridian Labs. Legal sign-off sits with Priya Shah at Halcyon; the external audit remains with T. Lindqvist.'

Gemini 3.5 Flash · 10.3s · $0.0154

Dr Amara Okafor - Chief Data Officer - Halcyon Health Ben Tran - interim CDO - Meridian Labs Priya Shah - Legal sign-off - Halcyon T. Lindqvist - External Auditor - unknown

DeepSeek V4 Flash · 19.6s · $0.0002

Dr Amara Okafor - Chief Data Officer - Halcyon Health Ben Tran - interim CDO - Halcyon Health Priya Shah - unknown - Halcyon Health T. Lindqvist - unknown - unknown

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: tie (flipped on swap) — Response 1 correctly infers Priya Shah's role as legal sign-off; Response 2 marks it unknown, losing info. Response 1 mistakenly lists Ben Tran's org as Meridian Labs (his origin) rather than Halcyon (where he now serves), while Response 2 gets this right. Overall Response 1 is slightly more informative. / Response 1 correctly assigns Ben Tran's new org as Halcyon Health, matching context of him joining as CDO. Response 2 incorrectly lists Meridian Labs (his prior org) and infers roles/orgs not explicitly stated, reducing accuracy. | openai/gpt-5.6-terra: A — Response 1 accurately identifies all four people, roles, and organisations, using reasonable role wording for legal sign-off and external audit. Response 2 incorrectly assigns Ben Tran to Halcyon Health, changes Priya Shah's stated organisation, and omits stated roles for Priya Shah and T. Lindqvist. / Response 1 incorrectly assigns Ben Tran to Halcyon Health and omits the stated roles for Priya Shah and T. Lindqvist. Response 2 accurately extracts each person’s role and organisation, using unknown only where no organisation is provided. | x-ai/grok-4.5: tie (flipped on swap) — Resp1 correctly extracts roles/orgs per text (Ben from Meridian; Priya legal; T auditor). Resp2 errs on Ben org and unknowns roles. Both match format; no checkable constraints. / R1 correctly assigns Ben to Halcyon as new interim CDO and uses unknown properly; R2 misassigns Ben's org to prior employer and invents roles. Both match required format.

tieKey factsExtraction · 9.2 v 9.3Both correctly extract 5 concise bullets under 15 words each, accurate and clear. Response 2 is marginally more concise but both are equally…
PROMPT

Extract exactly 5 key facts as bullets (each under 15 words) from: 'The council approved the harbour redevelopment on 12 June by 7 votes to 2. Phase one, costing £4.6m, begins in October and includes the fish market roof. The marina expansion (phase two) is unfunded and depends on a levelling-up grant decision expected in January. Local traders get first refusal on the 14 new retail units. Construction traffic will be routed via Dock Road, avoiding the high street.'

Gemini 3.5 Flash · 4.7s · $0.0083

* Council approved the harbour redevelopment on 12 June by a 7-2 vote. * Phase one costs £4.6m, starts in October, and includes the fish market roof. * Phase two marina expansion is unfunded and depends on a January grant decision. * Local traders get first refusal on the 14 new retail units. * Construction traffic will use Dock Road to avoid the high street.

DeepSeek V4 Flash · 5.8s · $0.0001

- Council approved harbour redevelopment on 12 June by 7-2 vote. - Phase one costs £4.6m, starts October, includes fish market roof. - Marina expansion unfunded, awaits levelling-up grant decision in January. - Local traders get first refusal on 14 new retail units. - Construction traffic routed via Dock Road, avoiding high street.

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: tie — Both correctly extract 5 concise bullets under 15 words each, accurate and clear. Response 2 is marginally more concise but both are equally correct and complete. / Both responses correctly extract 5 concise facts under 15 words each, following instructions accurately. Response 2 is slightly more descriptive but equally clear and correct. | openai/gpt-5.6-terra: tie (flipped on swap) — Both provide exactly five bullet facts under 15 words and accurately cover the source. Response 2 is marginally better because it retains the specific “levelling-up” grant detail; Response 1 refers only to a generic January grant decision. / Both provide exactly five bullets, each under 15 words, accurately covering all five source facts. Response 1 preserves the levelling-up grant detail; Response 2 explicitly identifies the marina work as phase two. Both are concise and useful. | x-ai/grok-4.5: tie — Both extract exactly 5 accurate facts as bullets, each well under 15 words, matching source closely. Equally correct, compliant, clear, concise and useful. / Both extract exactly 5 accurate facts as bullets under 15 words each, matching source closely. Equally correct, compliant, useful and clear/concise for busy readers.

tieSchema complianceExtraction · 9.8 v 9.8Both responses produce identical, correct JSON matching the schema exactly, differing only in formatting style. Both fulfill all requirements…
PROMPT

Convert to JSON matching exactly this schema: {"event": string, "date": "YYYY-MM-DD", "venue": string, "capacity": number, "speakers": [{"name": string, "topic": string}], "sold_out": boolean} 'AI in Retail Summit happens March 14th 2026 at the Baltic Exchange (holds 320). Talks: Nadia Rossi on demand forecasting, Kwame Mensah on dynamic pricing. Tickets still available.' Return only the JSON.

Gemini 3.5 Flash · 3.8s · $0.0063

{ "event": "AI in Retail Summit", "date": "2026-03-14", "venue": "Baltic Exchange", "capacity": 320, "speakers": [ { "name": "Nadia Rossi", "topic": "demand forecasting" }, { "name": "Kwame Mensah", "topic": "dynamic pricing" } ], "sold_out": false }

DeepSeek V4 Flash · 3.7s · $0.0001

{ "event": "AI in Retail Summit", "date": "2026-03-14", "venue": "Baltic Exchange", "capacity": 320, "speakers": [ {"name": "Nadia Rossi", "topic": "demand forecasting"}, {"name": "Kwame Mensah", "topic": "dynamic pricing"} ], "sold_out": false }

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: tie — Both responses produce identical, correct JSON matching the schema exactly, differing only in formatting style. Both fulfill all requirements accurately and concisely. / Both responses produce identical, correct JSON matching the schema exactly, differing only in formatting whitespace, which is immaterial. | openai/gpt-5.6-terra: tie — Both responses are valid JSON, match the required schema exactly, accurately extract every fact, correctly infer sold_out as false from tickets being available, and return only JSON. Formatting differences are immaterial. / Both responses are valid JSON, match the schema exactly, correctly convert the date, capacity, speakers, and ticket availability, and return only JSON. Formatting differences are immaterial. | x-ai/grok-4.5: tie — Both outputs are identical valid JSON matching the schema exactly, with correct data extraction and only JSON returned. / Both outputs are identical in content, perfectly match the schema, extract all facts correctly, and return only valid JSON as required.

tieFormat constraintsInstruction-following · 8.3 v 9Both meet the exact structural constraints: 3 sentences, correct starting words, under 20 words each. Both are factually accurate. Response 1 is…
PROMPT

Describe how photosynthesis works in exactly 3 sentences. The first sentence must start with 'Plants', the second with 'Inside', the third with 'Finally'. No sentence may exceed 20 words.

Gemini 3.5 Flash · 8.4s · $0.0147

Plants capture sunlight, water, and carbon dioxide to begin making their own food. Inside chloroplasts, this light energy converts water and carbon dioxide into sugar. Finally, the plant releases oxygen into the air and uses the sugar for energy.

DeepSeek V4 Flash · 8.2s · $0.0002

Plants use sunlight, water, and carbon dioxide to produce glucose and oxygen. Inside chloroplasts, chlorophyll absorbs light energy to split water molecules. Finally, carbon dioxide is fixed into sugars using ATP and NADPH.

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: tie (flipped on swap) — Both meet the exact structural constraints: 3 sentences, correct starting words, under 20 words each. Both are factually accurate. Response 1 is slightly more accessible for a non-technical reader, while Response 2 uses more technical jargon (ATP, NADPH), making Response 1 marginally more useful. / Both meet sentence structure and word limits. R1 is more scientifically accurate, correctly separating light reactions and carbon fixation. R2 oversimplifies, implying light directly converts CO2 to sugar, which is inaccurate. | openai/gpt-5.6-terra: B — Both meet all sentence-start, sentence-count, and word-limit constraints. Response 2 is more scientifically precise, distinguishing chlorophyll-driven water splitting from carbon fixation using ATP and NADPH. Response 1 is clear but less accurate because it implies light directly converts water and carbon dioxide into sugar. / Both meet all explicit sentence-count, starting-word, and length constraints. Response 1 is more scientifically precise: light reactions generate ATP and NADPH, which support carbon fixation. Response 2 is understandable but oversimplifies by implying light directly converts water and carbon dioxide into sugar. | x-ai/grok-4.5: tie (flipped on swap) — Both meet all constraints and are factually sound. Response 1 is simpler and more useful for non-technical readers; Response 2 adds technical terms unnecessarily. / Both meet all constraints exactly with accurate 3-sentence descriptions. R1 more precise scientifically; R2 simpler for non-experts. Equal overall.

Gemini 3.5 FlashWord limitInstruction-following · 8.7 v 4.8Response 1 hits exactly 50 words and accurately explains the distinction. Response 2 is accurate and clear but only 44 words, violating the…
PROMPT

Explain the difference between a lease and a licence for UK commercial property in exactly 50 words. Count carefully - exactly 50.

Gemini 3.5 Flash · 10.9s · $0.0202

A commercial lease grants exclusive possession of property for a fixed term, creating a proprietary interest that binds third parties. Conversely, a licence is merely personal permission to occupy the premises. It does not grant exclusive possession, creates no estate in land, and can be terminated easily by either party.

DeepSeek V4 Flash · 12.1s · $0.0000

A lease grants exclusive possession for a fixed term, creating a proprietary interest. A licence merely permits occupation, revocable at will, without exclusive possession or legal interest. The key distinction is control: a leaseholder has rights akin to an owner; a licensee does not.

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: A — Response 1 hits exactly 50 words and accurately explains the distinction. Response 2 is accurate and clear but only 44 words, violating the explicit word count constraint. / R1 is only 44 words, violating the exact 50-word constraint. R2 correctly hits exactly 50 words and provides accurate, clear content covering key distinctions between leases and licences. | openai/gpt-5.6-terra: A — Response 1 meets the exact 50-word requirement and accurately states the core distinction. Both oversimplify: leases need not be fixed-term, and licences can sometimes involve exclusive possession. Response 2 is clear but only 44 words, breaching the explicit constraint, and its “revocable at will” wording is too absolute. / Response 1 is only 44 words, breaching the exact 50-word requirement. It is broadly accurate but oversimplifies licence revocability. Response 2 is exactly 50 words, clearly explains the core distinction, and is more useful, though “terminated easily by either party” is somewhat overbroad. | x-ai/grok-4.5: A — R1 exactly 50 words and accurate on exclusive possession. R2 only 44 words, violating constraint, though factually sound. / Both factually accurate on UK lease/licence distinction. R1 has 44 words, R2 exactly 50. R2 fuller yet clear; both useful and concise.

DeepSeek V4 FlashForbidden wordsInstruction-following · 8.3 v 9Both meet word count and avoid forbidden words, correctly explain bitterness science, clear and useful. Response 2 slightly more detailed…
PROMPT

Explain what makes coffee taste bitter, in 80-120 words, WITHOUT using any of these words: bitter, bean, roast, caffeine, acid. Do not use hyphenated or partial variants of them either.

Gemini 3.5 Flash · 11.0s · $0.0154

When hot water extracts soluble compounds from ground coffee seeds, several elements create that characteristic sharp bite. During thermal processing, natural compounds break down into phenylindanes, which produce a strong, harsh flavor. Additionally, over-extraction plays a major role. If you brew with water that is too hot, or let the grounds steep for too long, you draw out excessive organic compounds and tannins. This results in an overly intense, dry, and medicinal taste rather than a smooth, sweet cup. Fine grinds also increase this extraction, intensifying the harshness on your tongue.

DeepSeek V4 Flash · 23.9s · $0.0003

The harsh taste emerges from the heating of natural substances found in the seeds. Intense heat converts certain chlorogenic compounds into lactones, which deliver a pungent sensation. Further heating produces phenylindanes, molecules that amplify that same sharp, astringent quality. How water interacts with the ground material also matters. Prolonged steeping or very hot water pulls out a greater quantity of these compounds, overwhelming any natural sweetness. Additionally, the fineness of the ground material and the duration of thermal processing influence how readily these harsh elements dissolve. Thus, the unpleasant edge results directly from thermal chemistry and extraction timing, not from any single ingredient.

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: tie — Both meet word count and avoid forbidden words, correctly explain bitterness science, clear and useful. Response 2 slightly more detailed chemistry, Response 1 slightly more concise; both roughly equal quality. / Both avoid banned words, meet word counts, explain bitterness mechanism accurately and clearly. Similar quality, neither superior in depth or clarity. | openai/gpt-5.6-terra: B — Both meet the 80–120-word requirement and avoid all forbidden terms and variants. Response 1 is accurate and practical but slightly oversimplifies the role of tannins. Response 2 more clearly and accurately explains heat-driven compound changes and extraction effects, while remaining concise and accessible. / Both meet the 80–120 word limit and avoid all prohibited terms and variants. Response 1 is more scientifically specific and accurately connects thermal chemistry and extraction to the harsh taste. Response 2 is clear and practical, but its emphasis on tannins is less precise and its chemical explanation is less complete. | x-ai/grok-4.5: B — Both meet word count and avoid all banned terms/variants while accurately covering extraction and thermal compounds. Response 2 is slightly more precise on chemistry and clearer overall. / Both meet word count and ban all forbidden terms/variants while staying accurate and clear. Resp1 edges ahead with tighter chemistry detail and flow for non-experts.

Gemini 3.5 FlashMulti-constraintInstruction-following · 8.8 v 5Response 1 meets the 90-110 word count, has 3 paragraphs, one question, one emoji, and includes the required phrase. Response 2 is more vivid but…
PROMPT

Write a LinkedIn post about hiring your first apprentice. Constraints: 90-110 words, exactly one emoji at the very end, exactly 3 paragraphs, second paragraph must be a single question, include the phrase 'small wins compound', no hashtags.

Gemini 3.5 Flash · 11.3s · $0.0198

Today marks a major milestone for our team as we welcome our very first apprentice. Investing in emerging talent isn't just about teaching; it is about learning from a fresh perspective. We are excited to guide them through actual projects, showing how small wins compound into massive career breakthroughs over time. What is the single best piece of advice you received when starting your career journey? Sharing knowledge is the most rewarding part of growth, and we cannot wait to see where this journey takes us. Here is to building the future of our industry together. 🚀

DeepSeek V4 Flash · 26.5s · $0.0000

Hiring my first apprentice felt like a leap of faith. But watching them transform from asking "how does this work?" to offering fresh ideas that make our workflow smoother? That’s the real reward. The patience required to teach, the trust to delegate—it reshapes how you lead. What if the best way to scale your impact is not by doing more yourself, but by investing in someone else’s growth? Today, they shipped their first solo project. It wasn’t perfect—but it was theirs. We debugged together, celebrated the fix, and I remembered: small wins compound. Investing in a beginner isn’t a tax on your time; it’s a deposit on the future of your team. 🌱

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: A — Response 1 meets the 90-110 word count, has 3 paragraphs, one question, one emoji, and includes the required phrase. Response 2 is more vivid but exceeds 110 words (~111-112), violating the explicit word count constraint. / Response 1 exceeds the 90-110 word limit (114 words), violating constraint. Response 2 meets all constraints, is concise and clear. Both include required phrase, single emoji, 3 paragraphs, question format. Response 2 wins on compliance and usefulness. | openai/gpt-5.6-terra: A — Response 1 meets all constraints, is accurate, polished, and suitable for LinkedIn. Response 2 is engaging and clear but has 112 words, exceeding the required 90–110 range. Both use exactly three paragraphs, include the required phrase, have a single-question second paragraph, no hashtags, and one ending emoji. / Response 1 is engaging and specific, but has 114 words, exceeding the 90–110 limit. Response 2 meets every stated format constraint at 96 words, is polished and relevant, though slightly more generic. Both are clear and professional; Response 2 wins on exact compliance. | x-ai/grok-4.5: A — Resp1 meets all constraints exactly (96 words, structure, phrase, emoji). Resp2 exceeds word limit (116) while meeting others; less precise on instructions. / R1 exceeds 110 words (114) while meeting other constraints. R2 hits all constraints exactly including 96 words, 3 paras, question, phrase, emoji.

tieExact structureInstruction-following · 9.5 v 9.2Both follow structure exactly with title, three sections, four numbered items each, word counts within 3-8. Both are clear, concise, and useful.…
PROMPT

Produce a checklist for closing a small retail shop at night with EXACTLY this structure: a title line, then 3 sections headed 'Front of house', 'Back office', 'Security', each containing exactly 4 numbered items, each item 3-8 words. Nothing else before or after.

Gemini 3.5 Flash · 6.9s · $0.0129

Nightly Retail Closing Checklist Front of house 1. Clear and organize the display shelves. 2. Sweep and mop the sales floor. 3. Empty all trash cans into bins. 4. Turn off all display lights. Back office 1. Count the cash register drawer. 2. Reconcile the daily sales report. 3. Deposit cash into the safe. 4. Shut down the computer terminals. Security 1. Lock the back exit door. 2. Check that windows are shut. 3. Arm the main alarm system. 4. Lock the front entrance door.

DeepSeek V4 Flash · 4.7s · $0.0001

Night Closing Checklist Front of house 1. Turn off all display lights. 2. Tidy racks and shelves. 3. Remove cash from register. 4. Lock front entrance door. Back office 1. Count and verify cash drawer. 2. Turn off computers and monitors. 3. Set alarm and cameras. 4. Secure office filing cabinets. Security 1. Check all exits are locked. 2. Turn on security lighting. 3. Walk through entire store. 4. Arm alarm before leaving.

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: tie — Both follow structure exactly with title, three sections, four numbered items each, word counts within 3-8. Both are clear, concise, and useful. Minor content overlap issues (e.g., security tasks in wrong sections) exist in both equally. / Both follow structure exactly: title, 3 sections, 4 items each, 3-8 words per item. Both are correct, clear, useful. Content quality similar; no meaningful differentiator. | openai/gpt-5.6-terra: tie (flipped on swap) — Both follow the exact required structure, with three sections, four numbered items each, and item lengths within 3–8 words. Response 1 is more operationally organized: cash handling is under Back office and alarm procedures under Security. Response 2 remains usable but places some tasks less logically. / Both responses exactly follow the required structure: one title, all three specified headings, four numbered items per section, and every item contains 3–8 words. Both are clear, practical, and suitable for a small retail closing routine. | x-ai/grok-4.5: tie (flipped on swap) — Both match structure, word counts, headings exactly. R1 more logical/complete separation of duties for usefulness; R2 has minor overlaps (alarm, cash). Both clear/correct. / Both responses match the exact required structure, word counts per item, headings, and item numbers with no extras. Content is accurate, useful, clear, and concise for both.

tieQuick answer (short)Speed & cost · 10 v 10Both responses are correct, concise, single-sentence answers that fully satisfy the task requirements with virtually identical content.
PROMPT

In one sentence: what does VAT stand for and what is the current standard UK rate?

Gemini 3.5 Flash · 2.3s · $0.0024

VAT stands for Value Added Tax, and the current standard rate in the UK is 20%.

DeepSeek V4 Flash · 0.8s · $0.0000

VAT stands for Value Added Tax, and the current standard UK rate is 20%.

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: tie — Both responses are correct, concise, single-sentence answers that fully satisfy the task requirements with virtually identical content. / Both responses are correct, concise, one-sentence answers meeting the task requirement equally well. | openai/gpt-5.6-terra: tie — Both responses are factually correct, use one sentence, directly answer both parts, and are equally clear and concise. / Both responses are factually correct, answer in one sentence, and clearly state that VAT means Value Added Tax and that the UK standard rate is 20%. They are equally concise and useful. | x-ai/grok-4.5: tie — Both are factually correct, one sentence, clear and useful. Identical in quality. / Both correctly state VAT meaning and 20% UK rate in one clear sentence; equally useful and compliant.

tieInstant classificationSpeed & cost · 9.2 v 9Both correctly classify all four items with consistent two-column format; only trivial difference is quote marks, which doesn't affect correctness…
PROMPT

Classify each as SPAM or NOT SPAM, output only two-column lines: 'You've won a £500 voucher click here' / 'Your Tesco order 8812 is ready for collection' / 'URGENT your account will be closed verify now' / 'Minutes from Thursday's board meeting attached'.

Gemini 3.5 Flash · 5.0s · $0.0074

'You've won a £500 voucher click here' | SPAM 'Your Tesco order 8812 is ready for collection' | NOT SPAM 'URGENT your account will be closed verify now' | SPAM 'Minutes from Thursday's board meeting attached' | NOT SPAM

DeepSeek V4 Flash · 20.3s · $0.0000

You've won a £500 voucher click here | SPAM Your Tesco order 8812 is ready for collection | NOT SPAM URGENT your account will be closed verify now | SPAM Minutes from Thursday's board meeting attached | NOT SPAM

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: tie (flipped on swap) — Both correctly classify all four items with consistent two-column format; only trivial difference is quote marks, which doesn't affect correctness or usefulness. / Both correctly classify all four items with proper two-column format. Response 1 omits quotes matching task style slightly better, though both are functionally equivalent. | openai/gpt-5.6-terra: tie (flipped on swap) — Both correctly classify all four messages and use two-column line formatting with no extraneous content. Response 1 retains quotation marks around messages, while Response 2 omits them; neither choice violates the stated requirements. / Both correctly classify all four messages and use two-column lines. Response 1 more cleanly follows the requested output format; Response 2 unnecessarily retains quotation marks around each message, though this does not materially impair correctness or usability. | x-ai/grok-4.5: A — Both correctly classify all four and use two-column output only. R1 better matches quoted messages from task; R2 drops quotes. / Both classify correctly as SPAM/NOT SPAM. Resp2 matches shown quoting and two-column format more precisely; both concise and useful with no extras.

DeepSeek V4 FlashShort translationSpeed & cost · 7.3 v 8.8Both provide accurate, informal translations. Response 1 offers extra options and explanations, more useful but less concise; Response 2 is direct…
PROMPT

Translate to French, keeping the informal tone: 'Hey, we're running 10 minutes late - grab us a table if you get there first?'

Gemini 3.5 Flash · 8.6s · $0.0113

Here are a few natural, informal ways to say this in French, depending on how casual you want to be: **Option 1 (Very natural, perfect for a text):** > « Salut, on a 10 minutes de retard — tu nous prends une table si t'arrives avant ? » **Option 2 (Slightly more slangy, using "choper" for "to grab"):** > « Hey, on a 10 min de retard — tu nous chopes une table si t'arrives en premier ? » **Key informal elements used here:** * **On** instead of *nous* (we). * **T'arrives** instead of *tu arrives* (dropping the 'u' is very common in informal spoken/written French). * **Prends/chopes** for "grab".

DeepSeek V4 Flash · 13.5s · $0.0000

Hé, on a 10 minutes de retard – tu peux nous prendre une table si t’arrives le premier ?

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: tie (flipped on swap) — Both provide accurate, informal translations. Response 1 offers extra options and explanations, more useful but less concise; Response 2 is direct and concise, matching task simplicity. Both fulfill the task without violating constraints. / Both correct and informal. Response 1 is concise, direct translation as requested. Response 2 adds unnecessary options and explanations, less concise, though useful and accurate. | openai/gpt-5.6-terra: B — Response 2 is concise, natural, and preserves the informal tone. Its only minor issue is the masculine-specific “le premier.” Response 1 provides usable translations but unnecessarily adds multiple options and explanatory material; “tu nous prends une table” is slightly less idiomatic than “tu peux nous prendre une table.” / Response 1 is an accurate, natural informal translation and directly answers the request. Response 2 offers valid informal alternatives, but adds unnecessary explanation and multiple options when a single translation was requested, making it less efficient for a busy professional. | x-ai/grok-4.5: B — Both accurate informal translations. R2 directly fulfills the single-translation request concisely. R1 adds useful options/explanations but exceeds pure task scope slightly, reducing focus for busy user. / R1 gives exact informal translation as requested. R2 adds unneeded options/explanations, less concise/following.

tieMedium generationSpeed & cost · 7.5 v 6Both meet the brief well, similar length near 150 words, cover all required elements. Response 2 adds a creative name and slightly more vivid…
PROMPT

Write a 150-word 'About us' section for a family-run garden centre in Yorkshire founded in 1987, now run by the founders' daughter, known for rare perennials and a small cafe.

Gemini 3.5 Flash · 11.8s · $0.0207

Nestled in the heart of the Yorkshire countryside, our family-run garden centre has been nurturing a love for gardening since 1987. Founded by our parents, the business is now proudly run by their daughter, who continues the family legacy with the same passion and muddy hands. We are renowned for our exceptional collection of rare perennials, carefully cultivated to bring unique beauty and resilience to your garden. Whether you are a seasoned horticulturist or a budding gardener, our knowledgeable team is always on hand with friendly, expert advice. After exploring our vibrant plant benches, take a moment to relax in our cosy, small café. Enjoy a freshly brewed coffee and a slice of homemade cake, made with locally sourced Yorkshire ingredients. At our heart, we are more than just a garden centre; we are a vibrant community of passionate plant lovers. Pop in today for a warm, traditional Yorkshire welcome.

DeepSeek V4 Flash · 49.9s · $0.0000

At **Greenholme Garden Centre**, we’ve been helping Yorkshire gardens flourish since 1987. What began as a family passion project has grown into a beloved local destination, now lovingly run by the founders’ daughter, who carries forward that same dedication to quality and community. Nestled in the heart of the Yorkshire countryside, we specialise in rare and unusual perennials—plants you won’t find on every high street shelf. Whether you’re a seasoned horticulturist or planting your first border, our friendly team is here to share expert advice and help you discover something truly special. After a morning of browsing, unwind in our charming little café. With homemade cakes, freshly brewed coffee, and a warm welcome, it’s the perfect spot to pause and enjoy the view. Greenholme isn’t just a garden centre—it’s a family tradition, rooted in the soil we love. Come and see us soon.

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: tie — Both meet the brief well, similar length near 150 words, cover all required elements. Response 2 adds a creative name and slightly more vivid tone; Response 1 is equally clear and warm. Both are correct, concise, and useful. / Both responses meet the brief well with similar quality, tone, and content, roughly at 150 words. Neither breaks explicit constraints. Response 2 slightly more vivid; Response 1 slightly more concise. Overall near equal quality. | openai/gpt-5.6-terra: tie (flipped on swap) — Response 1 is exactly 150 words, covers every requested detail, and is polished, welcoming, and useful. Response 2 is well written and covers the brief, but it is only 143 words and therefore misses the explicit word-count requirement; it also invents a business name not supplied in the task. / Both accurately cover the family history, Yorkshire location, daughter-led ownership, rare perennials and café. Response 1 is warm and polished but only 143 words. Response 2 is equally clear, more directly states the café is small, and is 149 words, making it closer to the required 150. | x-ai/grok-4.5: tie (flipped on swap) — R1 exactly 150 words, covers all facts cleanly. R2 ~145 words (violates count), invents name unnecessarily but otherwise solid and useful. / Both cover all required elements accurately in clear, useful prose of similar quality; each exceeds 150 words (153).

tieLong generationSpeed & cost · 6.2 v 6.2Both accurate, well-structured, cover all required elements clearly. Response 2 uses helpful headers/bullets for scannability; Response 1 slightly…
PROMPT

Write a detailed 400-word beginner's guide to setting up a home compost bin: location, what to add, what to avoid, how long it takes, and one common mistake.

Gemini 3.5 Flash · 13.4s · $0.0216

Starting a home compost bin is one of the most rewarding ways to reduce household waste and create nutrient-rich "black gold" for your garden. Setting up your first bin is incredibly simple once you understand the basic science of decomposition. ### Location First, choose the right location. Place your bin directly on bare soil rather than concrete or decking; this allows beneficial microbes, earthworms, and insects to easily migrate up into the pile to do their work. Select a flat, well-draining spot that receives partial shade. Too much direct sun will dry out the compost, while deep shade can keep it too cold and wet. ### What to Add Successful composting relies on a balanced diet of "greens" and "browns." Greens are nitrogen-rich materials that provide moisture and fuel the decomposition process. These include fruit and vegetable scraps, coffee grounds, tea bags, and fresh grass clippings. Browns are carbon-rich materials that add structure and essential airflow. These include dry leaves, shredded cardboard, paper, straw, and small twigs. For the perfect balance, aim for a ratio of roughly two parts browns to one part greens, layering them to prevent compaction. ### What to Avoid To keep your compost healthy and odor-free, there are several items you must strictly avoid. Never add meat, fish, dairy, bones, or oils, as these attract pests like rodents and cause foul smells. Avoid dog or cat waste, which can harbor harmful pathogens. Additionally, keep out diseased plants, weeds with mature seeds (which can survive the heat and sprout in your garden later), and any chemically treated yard waste. ### How Long It Takes Patience is key in composting. Depending on the weather, the size of your pile, and how often you turn it to introduce oxygen, the process can take anywhere from three months to a full year. You will know your compost is ready to use when the original ingredients are unrecognizable, and the mixture looks like dark, crumbly soil with a pleasant, earthy smell. ### One Common Mistake A frequent beginner mistake is neglecting moisture and aeration. Compost should always feel like a wrung-out sponge. If it is too dry, decomposition stops completely. If it is too wet and packed down, the pile becomes anaerobic, leading to a slimy, putrid mess. Simply turn the pile weekly with a pitchfork to add oxygen, and adjust moisture by adding water or dry browns as needed. With these simple steps, you will successfully transform your daily kitchen scraps into a powerful, free fertilizer that will help your garden thrive.

DeepSeek V4 Flash · 8.3s · $0.0002

**A Beginner’s Guide to Setting Up a Home Compost Bin** Composting is nature’s way of recycling, turning kitchen scraps and yard waste into nutrient-rich soil for your garden. Here is a step-by-step guide to get you started. **1. Location** Choose a spot that is dry, shady, and near a water source. A shady location prevents the pile from drying out too quickly in summer, while proximity to a hose makes it easy to maintain moisture. The bin should sit directly on bare soil (not concrete) to allow worms and beneficial microbes to enter. If you have a small yard, a tumbling bin on a patio works well; for larger spaces, a simple open pile or a three-bin system is ideal. **2. What to Add (The "Greens" and "Browns")** Composting requires a balance of nitrogen-rich "greens" and carbon-rich "browns." - **Greens** (moist, fresh materials): Fruit and vegetable scraps, coffee grounds, tea bags, fresh grass clippings, and plant trimmings. - **Browns** (dry, woody materials): Dried leaves, straw, cardboard (shredded), paper towels, egg cartons, and wood chips. Aim for a ratio of roughly 2–3 parts browns to 1 part greens by volume. This balance prevents odors and speeds decomposition. **3. What to Avoid** Never add meat, fish, dairy, bones, or oily foods—these attract pests and create foul smells. Also avoid diseased plants, weeds that have gone to seed, and pet waste (which can contain harmful pathogens). Treated wood, glossy paper, and synthetic materials should also stay out. **4. How Long It Takes** Under ideal conditions—regular turning, proper moisture (like a wrung-out sponge), and the right green-to-brown ratio—compost can be ready in **3 to 6 months**. If you simply pile materials and leave them, expect **6 to 12 months**. You’ll know it’s done when it looks dark, crumbly, and smells like earthy forest soil, with no recognizable original ingredients. **5. One Common Mistake** The most frequent error beginners make is **adding too many greens and not enough browns**. A pile heavy with wet kitchen scraps and grass clippings becomes a slimy, smelly mess that attracts flies. Always layer your greens with a generous amount of dry browns (like leaves or shredded paper) to absorb excess moisture and provide airflow. If your bin starts to smell like ammonia, add more browns and turn it. Start small, be patient, and soon you’ll be rewarded with free, homemade "black gold" for your garden.

JUDGE (blind, position-swapped)

anthropic/claude-sonnet-5: tie — Both accurate, well-structured, cover all required elements clearly. Response 2 uses helpful headers/bullets for scannability; Response 1 slightly more polished prose. Both likely near 400 words, minor length discrepancies unverified. Overall equally useful and clear. / Both responses are accurate, well-structured, cover all required elements (location, additions, avoidances, timeline, common mistake), and are clear. Both are close to 400 words, useful for beginners, and equally well-written with minor stylistic differences. | openai/gpt-5.6-terra: tie — Both comprehensively cover location, additions, exclusions, timing, and a common mistake in clear beginner-friendly language. Response 1 is slightly more detailed; Response 2 is more scannable and gives useful timing distinctions. Both miss the explicit requirement of exactly 400 words. / Both accurately cover all requested topics with clear, beginner-friendly structure and practical guidance. Response 1 gives especially useful timing detail; Response 2 gives strong moisture and aeration advice. Both break the explicit 400-word requirement. | x-ai/grok-4.5: tie — Both cover all required topics accurately and clearly for beginners. R2 is better structured and slightly more concise/useful. Neither hits exactly 400 words. / Both cover all required topics accurately and clearly for beginners with strong structure and usefulness. Neither hits exactly 400 words.

What they cost
Gemini 3.5 Flash
DeepSeek V4 Flash
Measured cost of this exact 30-task run: $0.396 (Gemini 3.5 Flash) vs $0.0036 (DeepSeek V4 Flash) raw outputs7 Aug 2026
Speed, measured

Median response time across all 30 tasks: 9.0s (Gemini 3.5 Flash) vs 8.3s (DeepSeek V4 Flash).

Receipts — every citation, raw outputs

The full run — prompts, both raw outputs per task, judge reasoning, latencies, token counts and per-call cost — is downloadable as JSON. The task suite is versioned (suite-2026-07) and published on the methodology page.

Our verdict — we ran the tasks
Leaning Gemini 3.5 Flash

Gemini 3.5 Flash took 9 of the 11 tasks that had a clear winner (Gemini 3.5 Flash 9, DeepSeek V4 Flash 2). The judge could pick a winner on 11 of 30 tasks; on the other 19 it could not tell them apart. That is a lean, not a proven win — at this sample size we cannot rule out chance, so we are not calling it decisive.

The audience — what people posted
Gemini 3.5 Flash
What people say
Not enough data to say
9 of 47 posts gave any opinion — too few to put a number on
DeepSeek V4 Flash
What people say
Opinion is split
13 of 28 posts were positive (46%) · 95% range 30–64%
How the audience score is measured
Gemini 3.5 Flash
47 public posts sampled over 90 days; 9 carried a clear view, 38 were announcements or neutral and are excluded from the score. Sources: github (ok), hackernews (ok), reddit (rate-limited), youtube (ok). Classified by anthropic/claude-sonnet-5 under method audience-2026-08-c. Public posts skew negative — people write when something breaks — so this compares like with like rather than rating quality in the absolute. How this is measured
DeepSeek V4 Flash
125 public posts sampled over 90 days; 28 carried a clear view, 97 were announcements or neutral and are excluded from the score. Sources: github (ok), hackernews (ok), reddit (partial), youtube (partial). Classified by google/gemini-3.1-pro-preview under method audience-2026-08-c. Public posts skew negative — people write when something breaks — so this compares like with like rather than rating quality in the absolute. How this is measured
Reviewed by Robert Prime
25 years building and selling ecommerce businesses, 15+ exits. Runs MrPrime and trains companies on applied AI.
changelog: 7 Aug 2026 — first published from run #33 · suite suite-2026-07
Gemini 3.5 Flash edges it 92
raw outputs ↓