Battles / Best free
Gemini 3.5 Flash wins 4–2
a narrow win on the tasks that separated them (6 of 30 tasks were decisive) — close enough that the loser is still worth a look.
Grok 4.5 vs Gemini 3.5 Flash · Best freemodels in this battle: Grok 4.5 vs Gemini 3.5 Flash
Show the full workings
The evidence
Suite-by-suite
blue = Grok 4.5 wins · grey = ties · white = Gemini 3.5 Flash wins (5 tasks per suite)
Round-by-round — all 30 tasks
tieCold emailWriting · 9.2 v 8.3Both meet constraints and are professional. Response 2 has a clearer, more specific CTA (5-min call Tuesday) and mentions brewery-specific pain…
Write a cold email (max 120 words) from the owner of a 6-person UK accounting firm to a local brewery, offering a free first-month bookkeeping trial. Friendly but professional, one clear call to action, no buzzwords.
**Subject:** Free first-month bookkeeping for [Brewery Name] Hi [Owner’s Name], I’m James Harris, owner of Harris & Co, a 6-person accounting firm based nearby. I’ve tried your beers and wanted to reach out directly. We handle straightforward bookkeeping for local businesses. I’d like to offer [Brewery Name] a free first month of bookkeeping so you can see how we work—no cost, no commitment. If that sounds useful, just reply to this email and we’ll set it up. Best regards, James Harris Owner, Harris & Co [phone] | [email] (Word count: 98)
Subject: Bookkeeping help for [Brewery Name] Hi [Name], I’m [My Name], and I run [Firm Name], a six-person accounting team based just down the road from you in [Town]. We’re big fans of your beers, but we also know that managing brewery finances—from stock tracking to VAT—can be a real headache. We help local independent businesses keep their books spotless so they can focus on what they do best. To show you how we can take some admin off your plate, I’d love to offer you a free first-month bookkeeping trial. No commitment, just a chance to see how we work. Would you be open to a quick 5-minute phone call next Tuesday? Best regards, [My Name] [Firm Name]
anthropic/claude-sonnet-5: tie (flipped on swap) — Both meet constraints and are professional. Response 2 has a clearer, more specific CTA (5-min call Tuesday) and mentions brewery-specific pain points, adding relevance and usefulness, while remaining concise. / R1 exceeds the 120-word limit (~135 words), violating the explicit constraint; R2 stays within limit (98 words) and is clear, professional, with a single CTA. Both are factually fine, but R2 better follows instructions and is more concise. | openai/gpt-5.6-terra: A — Both meet the word limit, tone, sender, offer, and single-CTA requirements. Response 1 is more concise and directly focused on the free trial. Response 2 is still strong but adds slightly generic phrasing and assumes brewery finance pain points, making it marginally less streamlined. / Both meet the 120-word limit, offer a free first month, use a friendly professional tone, avoid obvious buzzwords, and give one CTA. Response 2 is more direct and concise, clearly identifies the sender as owner, and makes replying exceptionally easy. Response 1 is strong but slightly more generic and its proposed call is marginally higher-friction. | deepseek/deepseek-v4-pro: tie (flipped on swap) — Response 1 is more concise (83 words) with a simpler reply CTA, perfectly meeting all constraints. Response 2 is slightly wordier and its phone-call CTA is less convenient for a busy professional. Both avoid buzzwords and stay under 120 words. / Response 1 personalizes with brewery-specific issues (VAT, stock) and asks a specific call, boosting usefulness. Response 2 is generic. Both meet constraints, but 1 better fits a busy professional.
Gemini 3.5 FlashProduct descriptionWriting · 8.3 v 8.9Both meet word count and avoid cliché, but Response 2 is more engaging, vivid, and tailored to gym-goers while remaining accurate and concise.…
Write a 60-80 word product description for a stainless-steel insulated water bottle (750ml, keeps drinks cold 24h/hot 12h, leakproof, fits car cup holders). Target audience: gym-goers. Avoid cliches like 'stay hydrated in style'.
This 750ml stainless-steel insulated water bottle is built for your gym routine. It keeps drinks ice-cold for 24 hours or hot for 12, so your water stays refreshing through intense workouts. The leakproof design means no spills in your gym bag, and it fits perfectly in car cup holders for easy transport. Durable and reliable, it's the perfect companion for every training session.
Crush your personal bests with a bottle built for the grind. This 750ml stainless-steel powerhouse keeps your water ice-cold for 24 hours or your post-workout brew hot for 12. Engineered with a 100% leakproof lid, you can toss it in your gym bag without a second thought. Its sleek profile fits perfectly in car cup holders for your commute. Durable, sweat-free, and ready for your toughest training sessions.
anthropic/claude-sonnet-5: B — Both meet word count and avoid cliché, but Response 2 is more engaging, vivid, and tailored to gym-goers while remaining accurate and concise. Response 1 is functional but generic. / Both meet word count and avoid the banned cliche, but Response 1 is more vivid and gym-specific, while Response 2 uses a mild cliche ('perfect companion') and is slightly less engaging. | openai/gpt-5.6-terra: tie (flipped on swap) — Both meet the 60–80 word requirement and include all key features for gym-goers. Response 1 is clearer, more direct, and avoids unsupported claims. Response 2 is engaging but adds unprovided “sweat-free” and “100%” claims, and its more slogan-heavy phrasing is less concise. / Both meet the 60–80-word requirement, target gym-goers, include all specified features, and avoid the cited cliché. Response 1 is more energetic and feature-rich but uses slightly promotional phrasing. Response 2 is clearer and more direct for a busy buyer, though more generic. Neither has a meaningful accuracy or instruction-following issue. | deepseek/deepseek-v4-pro: B — Response 2 uses more gym-specific, energetic language (e.g., 'crush your personal bests', 'sweat-free') while maintaining clarity and accuracy. Response 1 is clear but less targeted. Both meet all constraints. / Both are accurate, follow instructions, and target gym-goers. Response 1 uses more energetic, gym-specific language ('crush your personal bests', 'built for the grind'), making it slightly more engaging for the audience while remaining clear and concise.
tieSummarise messy notesWriting · 9.5 v 8.7Both are accurate 5-bullet summaries. Response 1 separates email campaign debate clearly, matching notes more precisely. Response 2 merges launch…
Turn these messy meeting notes into a clean 5-bullet summary a director could read in 20 seconds: 'ok so launch slipped again, marc says supplier issue?? maybe 3 wks. sarah wants to push the email campaign anyway but tom thinks wait. budget - we're 8k over, mostly the packaging redo. Q: do we tell retail partners now or after new date confirmed. also NEED to hire the warehouse temp before august rush. next mtg tues.'
- **Launch delay:** Slipped again (supplier issue per Marc); potential 3-week push. - **Email campaign:** Sarah wants to proceed now; Tom recommends waiting for new launch date. - **Budget:** $8k over, driven mainly by packaging redo. - **Retail partners:** Decide whether to notify now or after revised date is confirmed. - **Hiring + next steps:** Must hire warehouse temp before August rush; next meeting Tuesday.
Here is a 20-second summary for a director: * **Launch Delayed:** Slipped by ~3 weeks due to supplier issues; team is debating whether to delay the email campaign. * **Budget Deficit:** Currently $8k over budget, driven by the packaging redesign. * **Key Decision:** Determine whether to notify retail partners of the delay now or wait for a confirmed new date. * **Urgent Hiring:** Must onboard a warehouse temp before the August peak season begins. * **Next Meeting:** Tuesday (to finalize the timeline and partner communication plan).
anthropic/claude-sonnet-5: tie (flipped on swap) — Both are accurate 5-bullet summaries. Response 1 separates email campaign debate clearly, matching notes more precisely. Response 2 merges launch and campaign into one bullet, slightly less clean. Both concise and clear; Response 1 slightly more faithful to structure. / Both accurately capture all key points in 5 clean bullets, easy to scan in 20 seconds. Response 2 slightly better preserves the Sarah/Tom debate nuance; Response 1 slightly more polished formatting. Overall very comparable quality. | openai/gpt-5.6-terra: A — Response 1 is concise, exactly five bullets, and preserves uncertainty around the potential three-week delay. Response 2 is also clear and useful, but presents the delay more definitively and adds an unsupported claim that Tuesday’s meeting will finalize the timeline and communications plan. / Both provide clean, accurate five-bullet summaries covering all key issues. Response 2 is more concise and preserves attribution and uncertainty around the supplier issue and timing. Response 1 adds an unnecessary intro and slightly overstates Tuesday’s purpose as finalizing plans. | deepseek/deepseek-v4-pro: tie (flipped on swap) — Response 1 separates each point into a distinct bullet, making it easier to scan quickly, while Response 2 combines two points in the first bullet, slightly reducing clarity. Both are accurate and follow the 5-bullet constraint. / Both responses accurately capture all key points, use exactly 5 bullets, and are clear and concise for a director. No constraint violations.
Grok 4.5Tone rewriteWriting · 9.5 v 8.5Both responses are warm, accurate, and under 90 words. Response 1 is more concise while still covering all facts and tone requirements; Response 2…
Rewrite this complaint reply so it is warm, takes responsibility, and keeps the same facts, in under 90 words: 'Your order was delayed because of a system error on our side. The refund has been processed and will arrive in 3-5 days. We cannot offer further compensation. Let us know if you have questions.'
We’re truly sorry your order was delayed due to a system error on our side—that’s on us, and we take full responsibility. Your refund has already been processed and should arrive in 3–5 days. We’re unable to offer further compensation, but please don’t hesitate to reach out if you have any questions. We’re here to help.
I am so sorry your order was delayed. This was due to a system error on our side, and we take full responsibility for this mistake. To help make things right, we have processed your refund, which will arrive in your account within 3-5 business days. While we are unable to offer any further compensation, we truly appreciate your patience and understanding. Please let us know if you have any questions at all—we are here to help.
anthropic/claude-sonnet-5: A — Both responses are warm, accurate, and under 90 words. Response 1 is more concise while still covering all facts and tone requirements; Response 2 is slightly longer and more verbose but still compliant. / Both responses are warm, accurate, and within the word limit. Response 2 is more concise and equally complete, making it slightly more effective for a busy reader. | openai/gpt-5.6-terra: A — Both are warm, accountable, clear, and under 90 words. Response 1 preserves all facts exactly and is slightly more concise. Response 2 is also strong but changes “3–5 days” to “3–5 business days,” adding a detail not in the original. / Both are warm, accountable, accurate, and under 90 words. Response 2 is more concise and preserves the original timing exactly. Response 1 slightly changes “3–5 days” to “3–5 business days,” adding a detail not in the source. | deepseek/deepseek-v4-pro: abstain — unparseable
tieHeadlinesWriting · 8.8 v 9Both meet the mix requirement and word limits, are clear and useful. Response 2 explicitly labels word counts, aiding verification; Response 1 is…
Write 5 headline options (max 9 words each) for a blog post about how small restaurants can use AI to reduce food waste. Mix: 2 practical, 2 curiosity-driven, 1 with a number.
**Headline Options:** 1. **Practical:** Use AI Tools to Cut Restaurant Food Waste 2. **Practical:** Simple AI Steps for Less Kitchen Spoilage 3. **Curiosity-driven:** What If AI Knew Your Fridge Better? 4. **Curiosity-driven:** The Quiet Way AI Saves Tonight’s Specials 5. **With a number:** 5 AI Tricks Slash Small Restaurant Food Waste
Here are 5 headline options, all under the 9-word limit: **Practical** 1. **How Small Restaurants Can Cut Food Waste With AI** (9 words) 2. **Cut Restaurant Food Waste and Save Money With AI** (9 words) **Curiosity-Driven** 3. **Is AI the Secret to Zero Restaurant Food Waste?** (9 words) 4. **How AI is Quietly Solving Restaurant Food Waste** (8 words) **With a Number** 5. **5 Ways AI Helps Small Restaurants Cut Food Waste** (9 words)
anthropic/claude-sonnet-5: tie (flipped on swap) — Both meet the mix requirement and word limits, are clear and useful. Response 2 explicitly labels word counts, aiding verification; Response 1 is slightly more creative but less explicit. Both are correct and well-formatted, resulting in a near tie. / Both meet requirements with correct categories and word counts. Response 1's headlines are more directly relevant and clear; Response 2's curiosity headlines are more abstract/less clear on topic, slightly reducing usefulness. | openai/gpt-5.6-terra: tie — Both provide exactly five headlines, with two practical, two curiosity-driven, and one numbered option. All headlines are within nine words and relevant to AI food-waste reduction. Response 2 more consistently names small restaurants, while Response 1 is slightly punchier and avoids the potentially overstated “zero waste” claim. / Both provide exactly five headlines, meet the maximum nine-word limit, and include two practical, two curiosity-driven, and one numbered option. Response 1 is more explicitly targeted to small restaurants; Response 2 is slightly more concise and accessible. Neither has a meaningful accuracy or instruction-following issue. | deepseek/deepseek-v4-pro: tie (flipped on swap) — Both meet constraints. Response 1's headlines are more concise and varied, with stronger curiosity hooks like 'The Quiet Way AI Saves Tonight’s Specials'. Response 2 is slightly wordier and less punchy. Overall, 1 is more useful and clear. / Both meet all constraints. Response 1 explicitly mentions 'small restaurants' in every headline, making it slightly more targeted and useful for the intended audience.
tieBug fixCoding · 9.7 v 8.7Both correctly fix mutation and even-length median bugs with identical logic and a one-line explanation. Functionally equivalent, differing only…
This JavaScript function should return the median of a numeric array but gives wrong answers for even-length arrays and mutates the input. Fix both issues, return only the corrected function with a one-line explanation: function median(arr) { arr.sort(); const mid = Math.floor(arr.length / 2); return arr[mid]; }
```js function median(arr) { const sorted = [...arr].sort((a, b) => a - b); const mid = Math.floor(sorted.length / 2); return sorted.length % 2 ? sorted[mid] : (sorted[mid - 1] + sorted[mid]) / 2; } ``` Copies input to avoid mutation, numeric-sorts, and averages two middle values for even length.
```javascript function median(arr) { const sorted = [...arr].sort((a, b) => a - b); const mid = Math.floor(sorted.length / 2); return sorted.length % 2 !== 0 ? sorted[mid] : (sorted[mid - 1] + sorted[mid]) / 2; } ``` **Explanation:** This version clones the input array to prevent mutation, uses a numeric comparison function to sort the numbers correctly, and averages the two middle elements if the array has an even length.
anthropic/claude-sonnet-5: tie (flipped on swap) — Both correctly fix mutation and even-length median bugs with identical logic and a one-line explanation. Functionally equivalent, differing only in formatting; both fulfill task requirements equally well. / Both correct and fix both issues. Response 2 follows 'one-line explanation' instruction exactly, while Response 1's explanation spans multiple lines/sentences, slightly violating conciseness instruction. | openai/gpt-5.6-terra: tie (flipped on swap) — Both correctly clone the array, sort numerically, and average the middle pair for even lengths. Both provide a one-line explanation and satisfy the requested output structure. Response 1 is slightly more concise and direct; Response 2 is correct but more verbose than necessary. / Both correctly avoid input mutation via copying, use numeric sorting, and calculate the even-length median by averaging the two middle values. Each provides only the corrected function plus a one-line explanation. Response 2 is slightly more concise, but both fully satisfy the task. | deepseek/deepseek-v4-pro: tie (flipped on swap) — Both correctly fix mutation and even-length median, following the one-line explanation rule. Response 2's explanation is slightly clearer and free of typos, aiding a busy non-technical reader. / Response 1's explanation is multi-line, violating the one-line instruction. Response 2 is correct, concise, and follows all constraints exactly.
tieSmall featureCoding · 9.5 v 9.3Both implementations correctly handle batching, oversized invoices, and preserve order; both include exactly 3 valid asserts that pass. Response 1…
Write a Python function `chunk_invoices(invoices, max_total)` that takes a list of dicts like {'id': 'A1', 'amount': 120.5} and groups them into batches where each batch's summed amount does not exceed max_total. A single invoice larger than max_total goes in its own batch. Preserve order. Include 3 test cases using assert.
```python def chunk_invoices(invoices, max_total): """Group invoices into batches where each batch sum <= max_total. Invoices larger than max_total go in their own batch. Order is preserved. """ if not invoices: return [] batches = [] current_batch = [] current_sum = 0.0 for inv in invoices: amount = inv["amount"] if current_batch and current_sum + amount > max_total: batches.append(current_batch) current_batch = [] current_sum = 0.0 current_batch.append(inv) current_sum += amount if current_batch: batches.append(current_batch) return batches # Test 1: basic packing under the limit invoices1 = [ {"id": "A1", "amount": 100.0}, {"id": "A2", "amount": 50.0}, {"id": "A3", "amount": 80.0}, ] assert chunk_invoices(invoices1, 150.0) == [ [{"id": "A1", "amount": 100.0}, {"id": "A2", "amount": 50.0}], [{"id": "A3", "amount": 80.0}], ] # Test 2: single invoice larger than max_total gets its own batch invoices2 = [ {"id": "B1", "amount": 200.0}, {"id": "B2", "amount": 40.0}, {"id": "B3", "amount": 30.0}, ] assert chunk_invoices(invoices2, 150.0) == [ [{"id": "B1", "amount": 200.0}], [{"id": "B2", "amount": 40.0}, {"id": "B3", "amount": 30.0}], ] # Test 3: exact boundary + empty input invoices3 = [ {"id": "C1", "amount": 50.0}, {"id": "C2", "amount": 50.0}, {"id": "C3", "amount": 50.0}, ] assert chunk_invoices(invoices3, 100.0) == [ [{"id": "C1", "amount": 50.0}, {"id": "C2", "amount": 50.0}], [{"id": "C3", "amount": 50.0}], ] assert chunk_invoices([], 100.0) == [] ```
Here is the Python function `chunk_invoices` along with 3 test cases using `assert` to verify its correctness. ```python def chunk_invoices(invoices, max_total): """ Groups invoices into batches where each batch's total amount does not exceed max_total. An invoice with an amount greater than max_total is placed in its own batch. Preserves the original order of the invoices. """ batches = [] current_batch = [] current_sum = 0.0 for invoice in invoices: amount = invoice['amount'] # If a single invoice exceeds the max_total, it gets its own batch if amount > max_total: if current_batch: batches.append(current_batch) current_batch = [] current_sum = 0.0 batches.append([invoice]) # If adding the invoice keeps the batch under or equal to max_total elif current_sum + amount <= max_total: current_batch.append(invoice) current_sum += amount # If adding the invoice would exceed max_total, start a new batch else: batches.append(current_batch) current_batch = [invoice] current_sum = amount # Append the final batch if it contains any invoices if current_batch: batches.append(current_batch) return batches # ========================================== # Test Cases # ========================================== # Test Case 1: Standard grouping within limit invoices_1 = [ {'id': 'A1', 'amount': 50.0}, {'id': 'A2', 'amount': 30.0}, {'id': 'A3', 'amount': 40.0}, {'id': 'A4', 'amount': 10.0} ] # Batch 1: A1 + A2 = 80.0 (fits exactly) # Batch 2: A3 + A4 = 50.0 (fits) assert chunk_invoices(invoices_1, max_total=80) == [ [{'id': 'A1', 'amount': 50.0}, {'id': 'A2', 'amount': 30.0}], [{'id': 'A3', 'amount': 40.0}, {'id': 'A4', 'amount': 10.0}] ] # Test Case 2: Handling an invoice larger than max_total invoices_2 = [ {'id': 'B1', 'amount': 30.0}, {'id': 'B2', 'amount': 150.0}, # Exceeds max_total of 100 {'id': 'B3', 'amount': 40.0} ] # Batch 1: B1 (30.0) # Batch 2: B2 (150.0) - isolated because it exceeds max_total # Batch 3: B3 (40.0) assert chunk_invoices(invoices_2, max_total=100) == [ [{'id': 'B1', 'amount': 30.0}], [{'id': 'B2', 'amount': 150.0}], [{'id': 'B3', 'amount': 40.0}] ] # Test Case 3: All invoices fit into a single batch invoices_3 = [ {'id': 'C1', 'amount': 10.5}, {'id': 'C2', 'amount': 20.0}, {'id': 'C3', 'amount': 15.0} ] # Batch 1: C1 + C2 + C3 = 45.5 (fits under 50) assert chunk_invoices(invoices_3, max_total=50) == [ [{'id': 'C1', 'amount': 10.5}, {'id': 'C2', 'amount': 20.0}, {'id': 'C3', 'amount': 15.0}] ] print("All test cases passed successfully!") ```
anthropic/claude-sonnet-5: tie (flipped on swap) — Both implementations correctly handle batching, oversized invoices, and preserve order; both include exactly 3 valid asserts that pass. Response 1 is slightly more concise; Response 2 has extra commentary but is equally correct and clear. / Both correct and meet constraints. Response 2 is more concise, handles empty input edge case, and has cleaner logic without redundant special-casing for oversized invoices. | openai/gpt-5.6-terra: tie (flipped on swap) — Both implementations preserve order, correctly form contiguous batches, and include three labeled test cases with assertions. Response 2 explicitly isolates oversized invoices immediately, making its behavior clearer and more robust. Response 1 is concise and correct for normal nonnegative invoice amounts, but its oversized-invoice handling is less explicit. / Both implementations correctly preserve order, create sequential batches within the limit, and isolate oversized invoices. Each includes three labeled test cases using assertions. Response 2 additionally tests empty input; this does not conflict with the requirement. Both are clear and useful. | deepseek/deepseek-v4-pro: tie (flipped on swap) — Both functions correctly group invoices by sum, handle large invoices, preserve order, and include three valid assert tests. They are clear and useful. / Both are correct and follow all instructions. Response 2 handles large invoices implicitly and includes an empty-input test, making it slightly more concise and comprehensive. Response 1 is slightly more verbose but equally correct.
tieExplain codeCoding · 9.3 v 9Both correctly explain memoization and memory leak risk. Response 1 stays under 150 words (98) with confirmed count; Response 2 exceeds ~140 words…
Explain to a junior developer, in under 150 words, what this code does and one risk of using it: const cache = new Map(); function memo(fn) { return (...args) => { const k = JSON.stringify(args); if (!cache.has(k)) cache.set(k, fn(...args)); return cache.get(k); }; }
**This is a simple memoization helper.** `memo` takes any function `fn` and returns a wrapped version. The wrapper turns the arguments into a string key via `JSON.stringify`, checks a shared `Map` cache, and only calls `fn` on a cache miss. Later calls with the same args get the stored result instantly (handy for expensive pure calculations). **One risk:** The `cache` Map is module-scoped and never cleared or size-limited, so it grows forever. Long-running apps can leak memory, especially with many unique argument combinations. (Word count: 98)
This code is a **memoization** helper. It speeds up expensive functions by caching (saving) their results. When you wrap a function with `memo`, it converts the input arguments into a string key using `JSON.stringify`. If the wrapped function is called with those same arguments again, it skips the actual calculation and instantly returns the saved result from the `cache` Map. ### The Risk: Memory Leaks The global `cache` Map **never clears its data**. If your application runs continuously and is called with many unique arguments, the cache will grow indefinitely. This consumes more and more RAM over time, which can eventually crash your application. To use this safely in production, you would need a way to limit the cache size (like a Least Recently Used (LRU) eviction policy) or clear it periodically.
anthropic/claude-sonnet-5: tie (flipped on swap) — Both correctly explain memoization and memory leak risk. Response 1 stays under 150 words (98) with confirmed count; Response 2 exceeds ~140 words but doesn't state count, likely close to or over limit. Response 1 is more concise and equally accurate. / Both accurately explain memoization and identify the memory-leak risk within word limits. Similar clarity and usefulness; Response 2 slightly more concise, Response 1 slightly more detailed with actionable fix. Roughly equal quality. | openai/gpt-5.6-terra: A — Both accurately explain memoization and the unbounded shared cache memory-growth risk, and both are under 150 words. Response 1 is more concise and directly tailored to the requested explanation plus one risk. Response 2 is also clear but adds less necessary production advice. / Both accurately explain memoization and identify unbounded cache growth as a memory risk, while staying under 150 words. Response 2 is more concise and directly useful for a junior developer. Response 1 is also strong but adds slightly more detail than necessary. | deepseek/deepseek-v4-pro: tie (flipped on swap) — Both correctly explain memoization and the unbounded cache memory leak. Response 2 uses plainer language, slightly more useful for a non‑technical reader, but both are clear and within the word limit. Tie at 9. / Both are correct, clear, and under limit. Response 1 explains the risk more thoroughly, noting it's a global cache and suggesting an LRU eviction policy, which is more actionable for a junior dev. Response 2 is concise but slightly less helpful. Scores weighted accordingly.
tieSQL queryCoding · 8.8 v 8.8Both queries are correct; RANK() is arguably more standard for 'top 3' than DENSE_RANK. Response 1 is concise, meeting the single-query request…
Given tables orders(id, customer_id, created_at, total) and customers(id, name, country), write a single SQL query returning each country's top 3 customers by lifetime spend in 2025, with columns country, name, total_spend, rank. Use a window function. Standard PostgreSQL.
```sql SELECT country, name, total_spend, rank FROM ( SELECT c.country, c.name, SUM(o.total) AS total_spend, RANK() OVER (PARTITION BY c.country ORDER BY SUM(o.total) DESC) AS rank FROM customers c JOIN orders o ON o.customer_id = c.id WHERE o.created_at >= '2025-01-01' AND o.created_at < '2026-01-01' GROUP BY c.id, c.country, c.name ) t WHERE rank <= 3 ORDER BY country, rank; ```
Here is the standard PostgreSQL query to achieve this. It uses a Common Table Expression (CTE) to first calculate the total spend per customer for the year 2025, applies the `DENSE_RANK()` window function partitioned by country, and then filters for the top 3 in the outer query. ```sql WITH customer_spend_2025 AS ( SELECT c.country, c.name, SUM(o.total) AS total_spend, DENSE_RANK() OVER ( PARTITION BY c.country ORDER BY SUM(o.total) DESC ) AS rank FROM customers c JOIN orders o ON c.id = o.customer_id WHERE o.created_at >= '2025-01-01' AND o.created_at < '2026-01-01' GROUP BY c.id, c.country, c.name ) SELECT country, name, total_spend, rank FROM customer_spend_2025 WHERE rank <= 3 ORDER BY country, rank; ``` ### Explanation: 1. **Date Filtering**: `o.created_at >= '2025-01-01' AND o.created_at < '2026-01-01'` efficiently captures all orders placed in the year 2025, utilizing indexes on `created_at` if they exist. 2. **Aggregation**: `SUM(o.total)` calculates the lifetime spend for each customer within that 2025 timeframe. 3. **Window Function**: `DENSE_RANK() OVER (PARTITION BY c.country ORDER BY SUM(o.total) DESC)` ranks customers within their respective countries. `DENSE_RANK` is used so that if there is a tie for 3rd place, all tied customers are included. 4. **Outer Filter**: `WHERE rank <= 3` restricts the final output to only the top 3 customers per country.
anthropic/claude-sonnet-5: tie (flipped on swap) — Both queries are correct; RANK() is arguably more standard for 'top 3' than DENSE_RANK. Response 1 is concise, meeting the single-query request directly. Response 2 adds extra explanation not required, less concise. / Both queries are correct and functionally similar, using window functions correctly. Response 1 provides clearer explanation, aiding a non-technical reader, while Response 2 is more concise but lacks explanation. Response 1 slightly better overall for clarity and usefulness. | openai/gpt-5.6-terra: A — Both are valid PostgreSQL single-query solutions using window functions and correct 2025 date filtering. Response 1 uses RANK(), which better matches the requested rank column. Response 2's DENSE_RANK() changes rank values after ties and may return customers beyond the top three positions; its explanation of tie handling is imprecise. / Both queries correctly aggregate 2025 orders per customer, rank within each country using a window function, and filter to ranks 1–3. Response 2 is more concise and directly satisfies the request. Response 1 is also valid, though more verbose and DENSE_RANK has slightly different tie behavior than RANK. | deepseek/deepseek-v4-pro: B — Both queries correctly return top 3 customers per country by 2025 spend using window functions. Response 2 adds a clear explanation and uses DENSE_RANK to include ties, making it more useful for a non-technical professional. / Response 1 provides thorough explanation and uses DENSE_RANK with tie-handling, aiding understanding. Response 2 is correct but minimal, less useful for a non-technical professional.
tieRegexCoding · 8.7 v 8.3Both provide correct, near-identical regexes and valid JS one-liners with a reasonable edge case noted. Response 2's edge case is slightly more…
Write a regex that matches UK postcodes like 'SW1A 1AA', 'M1 1AE', 'B33 8TH' (allow lowercase and optional space), and a one-line JavaScript example using it to validate a form field. Briefly note one edge case your regex does NOT handle.
### Regex ```regex /^[A-Z]{1,2}\d[A-Z\d]? ?\d[A-Z]{2}$/i ``` ### One-line JS example ```js const valid = /^[A-Z]{1,2}\d[A-Z\d]? ?\d[A-Z]{2}$/i.test(document.getElementById('postcode').value); ``` ### Edge case not handled Does not match special postcodes like `GIR 0AA` (or BFPO/overseas formats).
### The Regex ```regex /^[A-Z]{1,2}[0-9][A-Z0-9]?\s?[0-9][A-Z]{2}$/i ``` ### JavaScript One-Line Validation Example ```javascript const isValid = /^[A-Z]{1,2}[0-9][A-Z0-9]?\s?[0-9][A-Z]{2}$/i.test(document.getElementById('postcode').value.trim()); ``` ### Edge Case Not Handled This regex **does not validate UK-specific letter restrictions**. To prevent confusion with numbers, the UK Royal Mail does not allow certain letters in specific positions. For example, the letters **I, L, and O** are never used in the last two characters of a postcode. This simplified regex would incorrectly mark `SW1A 1IO` as valid, even though it is physically impossible.
anthropic/claude-sonnet-5: tie (flipped on swap) — Both provide correct, near-identical regexes and valid JS one-liners with a reasonable edge case noted. Response 2's edge case is slightly more insightful; both meet instructions equally well. / Both meet requirements. Response 2's edge case (GIR 0AA special postcode) is more accurate and commonly cited than Response 1's letter-restriction claim, which is less standard. Response 2 is also more concise. | openai/gpt-5.6-terra: A — Response 1 correctly matches the examples, supports lowercase and an optional space, provides a true unsupported edge case, and gives a one-line JS use. Response 2 is broadly functional, but its letter-restriction explanation is factually inaccurate: L is not universally prohibited in the final two postcode characters. Its \s also permits non-space whitespace. / Both provide a case-insensitive regex, optional space, matching examples, one-line JavaScript validation, and an edge case. Response 2 accurately identifies GIR 0AA as unsupported. Response 1’s general limitation is valid, but its claim that L is never used in the final two characters is inaccurate; L can occur there. | deepseek/deepseek-v4-pro: B — Both regexes are correct and follow instructions. Response 2 adds .trim() for robustness and notes letter restrictions, making it slightly more useful. Response 1's edge case is valid but less detailed. / Both regex and notes are correct. Response 1's JS example trims input (more robust for forms) and its edge-case explanation is clearer, making it slightly more useful and concise for a non-technical user.
tieMulti-step logicReasoning · 9 v 9.3Both correct and identical calculations. Response 2 adds a helpful comparison table and caveats, slightly more useful for a non-technical reader…
A bakery sells loaves at £3.20. Ingredients cost £1.10/loaf, labour £0.90/loaf, fixed costs £480/week. They sell 450 loaves/week. A supplier offers ingredients at £0.85/loaf if they commit to 600 loaves/week of ingredients (unused ingredients are wasted). Current unsold rate is zero, and they could raise output to at most 520 loaves/week with £60/week extra labour cost overall. Should they take the deal? Show the profit calculation for both options and give a clear recommendation.
**Recommendation: take the deal.** It raises weekly profit from £60 to £146. ### Current weekly profit (no deal) - Revenue: \(450 \times 3.20 = 1440\) - Ingredients: \(450 \times 1.10 = 495\) - Labour: \(450 \times 0.90 = 405\) - Fixed costs: \(480\) - Total costs: \(495 + 405 + 480 = 1380\) - **Profit: \(1440 - 1380 = 60\)** ### Deal weekly profit (commit to 600 loaves of ingredients, produce/sell maximum 520) - Revenue: \(520 \times 3.20 = 1664\) - Ingredients (600 committed, 80 wasted): \(600 \times 0.85 = 510\) - Labour: \(520 \times 0.90 + 60 = 528\) - Fixed costs: \(480\) - Total costs: \(510 + 528 + 480 = 1518\) - **Profit: \(1664 - 1518 = 146\)** The deal is profitable only if output is raised to the new maximum of 520 (producing/selling only 450 under the deal yields a lower profit of £45). Marginal revenue on the extra loaves exceeds the extra labour cost and covers the cost of wasted ingredients, so accept the supplier offer.
Based on the financial calculations below, the bakery **should take the deal**, provided they can sell the increased output of 520 loaves per week. Here is the detailed profit breakdown for both options: --- ### Option A: Current Situation (450 loaves/week) In this scenario, the bakery sells all 450 loaves they produce. * **Weekly Revenue:** 450 loaves × £3.20 = **£1,440.00** * **Weekly Costs:** * *Ingredients:* 450 loaves × £1.10 = £495.00 * *Labour:* 450 loaves × £0.90 = £405.00 * *Fixed Costs:* £480.00 * *Total Costs:* £495 + £405 + £480 = **£1,380.00** * **Weekly Profit:** £1,440.00 − £1,380.00 = **£60.00** --- ### Option B: Proposed Deal (Maximized Output of 520 loaves/week) To make the deal viable, the bakery must increase production to its maximum capacity of 520 loaves. They must pay for 600 loaves' worth of ingredients, meaning 80 loaves' worth of ingredients will be wasted. * **Weekly Revenue:** 520 loaves × £3.20 = **£1,664.00** * **Weekly Costs:** * *Ingredients:* 600 loaves × £0.85 (committed amount) = £510.00 * *Labour:* (520 loaves × £0.90) + £60.00 extra = £468.00 + £60.00 = £528.00 * *Fixed Costs:* £480.00 * *Total Costs:* £510 + £528 + £480 = **£1,518.00** * **Weekly Profit:** £1,664.00 − £1,518.00 = **£146.00** *(Note: If they took the deal but kept production at 450 loaves, profit would actually drop to **£45.00** due to the ingredient waste).* --- ### Comparison Summary | Metric | Option A (Current) | Option B (New Deal) | Difference | | :--- | :--- | :--- | :--- | | **Loaves Sold** | 450 | 520 | +70 | | **Revenue** | £1,440.00 | £1,664.00 | +£224.00 | | **Total Costs** | £1,380.00 | £1,518.00 | +£138.00 | | **Weekly Profit** | **£60.00** | **£146.00** | **+£86.00** | --- ### Recommendation **Yes, take the deal.** By accepting the supplier's offer and increasing weekly production to 520 loaves, the bakery will **more than double its weekly profit** from £60 to £146 (an increase of 143%). **Key Considerations:** 1. **Market Demand:** This recommendation assumes the bakery can easily sell the extra 70 loaves per week at the current price of £3.20. 2. **Wastage:** Even though 80 loaves' worth of ingredients are wasted each week, the bulk discount on the remaining 520 loaves more than compensates for this loss.
anthropic/claude-sonnet-5: tie (flipped on swap) — Both correct and identical calculations. Response 2 adds a helpful comparison table and caveats, slightly more useful for a non-technical reader, though slightly longer. Clarity and thoroughness favor Response 2 marginally. / Both correctly compute identical figures and recommend taking the deal. Response 1 is more detailed with a comparison table, aiding a non-technical reader, but slightly longer. Response 2 is concise and equally accurate. Both meet task requirements well. | openai/gpt-5.6-terra: tie (flipped on swap) — Both calculations are correct: current profit is £60 and deal profit at 520 sold is £146, versus £45 if output remains 450. Response 2 is stronger because it explicitly conditions the recommendation on demand for the additional 70 loaves, which is not guaranteed by the prompt. Response 1 is more concise but states an unconditional recommendation. / Both calculate £60 current profit and £146 at 520 correctly, including 600 ingredients, £60 extra labour, and wastage. Both appropriately note the result depends on selling additional loaves. However, neither should imply 520 is strictly required: profitability exceeds £60 above roughly 483 loaves (assuming the £60 applies). | deepseek/deepseek-v4-pro: tie (flipped on swap) — Response 2 is equally accurate but more structured with a comparison table and explicit assumptions, making it clearer and more useful for a busy professional. / Both responses accurately compute current and deal profits, include the 450-loaf scenario under the deal, and offer clear recommendations. No factual errors or instruction breaches.
tiePlanningReasoning · 8.2 v 7.2Both meet format with owner per day and risk flag. Response 1 is clearer and more concise; Response 2 is solid but slightly more complex with less…
Plan the launch week for a small online course (already recorded). Resources: one founder, a part-time VA (10h), email list of 2,000, £300 ad budget. Produce a 7-day plan, one line per day, each line naming the owner. Flag the single riskiest dependency.
**7-Day Launch Plan** - **Day 1**: Founder finalizes sales page, pricing, and email sequence; VA sets up basic tracking/links. - **Day 2**: Founder writes/sends warm-up email to list; VA schedules remaining emails and preps simple creatives. - **Day 3**: Founder goes live or posts value content; VA launches £100 ads (retargeting + lookalikes) and monitors. - **Day 4**: Founder sends cart-open email + personal outreach; VA handles replies and basic support. - **Day 5**: Founder hosts Q&A or scarcity push; VA optimizes ads with remaining budget and tracks conversions. - **Day 6**: Founder sends deadline email; VA manages inbox, testimonials, and last ad tweaks. - **Day 7**: Founder closes cart + sends final email; VA wraps reporting and customer onboarding. **Riskiest dependency**: Email list engagement/deliverability (the 2,000-person list is the primary free channel; weak opens or spam issues kill momentum faster than ads or VA hours can fix).
**Monday (Founder):** Launch the cart by sending Email #1 (Cart Open) and activating the £300 Meta retargeting ad campaign. **Tuesday (VA):** Schedule Email #2 (Case Study/Value) and monitor/moderate comments on the active retargeting ads. **Wednesday (Founder):** Host a live Q&A on social media to address buyer objections and send Email #3 (FAQ). **Thursday (VA):** Update the sales page with a 24-hour fast-action bonus and schedule Email #4 (Bonus Announcement). **Friday (Founder):** Send Email #5 (Bonus Expiring) and optimize the retargeting ad budget toward high-intent custom audiences. **Saturday (VA):** Clear the customer support inbox of technical/billing questions and schedule Email #6 (Cart Closes Tomorrow). **Sunday (Founder):** Send two final urgency emails (morning/evening) and close the cart and ad campaigns at midnight. *** **Riskiest Dependency:** **Email deliverability and checkout tech integration.** With a small list of 2,000 and a tiny ad budget, 95%+ of sales will come directly from email links; if the payment gateway fails or emails land in spam on Day 1, the launch will fail.
anthropic/claude-sonnet-5: tie (flipped on swap) — Both meet format with owner per day and risk flag. Response 1 is clearer and more concise; Response 2 is solid but slightly more complex with less clear VA hour constraints (10h total). / Both follow format and flag same risk; Response 1 gives more concrete, specific daily actions tied to launch mechanics, slightly more useful and clear than Response 2's vaguer descriptions. | openai/gpt-5.6-terra: A — Response 1 gives a practical seven-day sequence, names responsible owners daily, respects the budget, and identifies one clear dependency. Response 2 is actionable but makes an unsupported “95%” claim and flags two dependencies (deliverability and checkout integration), despite requiring a single riskiest dependency. / Response 1 provides a more coherent full-week cart-open sequence, clear ownership, and practical use of email and ads. However, it identifies two riskiest dependencies rather than one. Response 2 includes useful preparation but delays cart opening until Day 4 and spreads a very small ad budget across retargeting and lookalikes. | deepseek/deepseek-v4-pro: tie (flipped on swap) — Response 1 follows all instructions exactly, naming a single riskiest dependency. Response 2 flags two dependencies, slightly missing the 'single' requirement. Both are clear and useful, but Response 1's strict adherence gives it the edge. / Response 1 gives a precise daily schedule with specific emails and ad activation, better for quick execution. Response 2 is good but somewhat vaguer. Both correctly flag email deliverability. Concise, clear, and meets all constraints.
tieTrade-off analysisReasoning · 9 v 7.8Both correctly recommend B, cover cost/flexibility/quality/non-obvious factor, stay under 200 words, clear formatting. Response 2 slightly more…
A 12-person agency must choose: (A) hire a mid-level developer at £55k, or (B) contract overflow work to freelancers at roughly £400/day, expected 60 days/year. Give a recommendation in under 200 words covering cost, flexibility, quality risk, and one non-obvious factor.
**Recommend (B) freelancers** for now. **Cost:** Freelancers total ~£24k/year (400 × 60). A hire costs £55k salary plus ~25-40% on-costs (NI, pension, benefits, equipment, recruitment), easily £70k+. Clear savings on pure overflow. **Flexibility:** Contractors scale to actual days with no idle salary, notice periods, or severance. Ideal for variable 60-day demand. **Quality risk:** Higher with freelancers (variable skill, context loss, less accountability). A mid-level hire offers consistency, code ownership, and team integration—but only if you can fill ~200+ days of productive work. **Non-obvious factor:** Knowledge retention and ramp-up tax. Every freelancer re-learns your stack, clients, and standards; a hire compounds familiarity and can mentor or productise internal tools, reducing future overflow. If overflow stays ≤60-80 days and is non-core, stick with (B) and a tight vetted bench. If demand is rising or strategic, switch to (A). Revisit in 6 months with utilisation data.
Recommend **Option B (Freelancers)**. * **Cost:** Option B costs £24,000/year. Option A costs £55,000 base, but rises to ~£71,500 after UK National Insurance, pension, and equipment. Freelancing saves the agency over £47,000 annually. * **Flexibility:** High. A 12-person agency faces volatile pipeline risk. Option B scales to £0 during dry spells, whereas Option A is a heavy, permanent fixed cost. * **Quality Risk:** High. Freelancers present inconsistent coding standards and onboarding friction. This must be mitigated by having your internal team conduct strict pull-request reviews. * **Non-Obvious Factor (Knowledge Leakage):** When freelancers leave, they take the undocumented "how and why" of the codebase with them. This creates a long-term maintenance liability for your agency, which remains contractually obligated to support the client. * **Verdict:** Choose Option B to protect cash flow. Transition to Option A only when overflow consistently exceeds 140 days/year (the financial tipping point where hiring becomes cheaper).
anthropic/claude-sonnet-5: tie — Both correctly recommend B, cover cost/flexibility/quality/non-obvious factor, stay under 200 words, clear formatting. Response 2 slightly more precise with tipping point calculation; Response 1 slightly more nuanced advice. Both equally strong overall. / Both correctly recommend B with cost math, flexibility, quality risk, and knowledge-retention as non-obvious factor. Both under 200 words, well-structured. Response 2 slightly more concise and adds actionable revisit trigger; Response 1 gives specific tipping point. Very comparable quality. | openai/gpt-5.6-terra: A — Response 1 accurately compares likely fully loaded employment cost with £24k freelance spend, covers all requested factors, and gives a practical decision rule. Response 2 is clear and useful, but its stated 140-day financial tipping point conflicts with its own ~£71.5k fully loaded cost; at £400/day it is closer to 179 days. / Both address all required factors and stay under 200 words. Response 2 is more balanced, practical, and gives a sensible decision rule and review point. Response 1’s stated 140-day financial tipping point is not supported by its own £71.5k fully loaded cost estimate; at £400/day it is closer to 179 days. | deepseek/deepseek-v4-pro: tie (flipped on swap) — Both responses accurately recommend Option B, cover cost, flexibility, quality risk, and a non-obvious factor, and stay under 200 words. They are equally correct, clear, and useful for a non-technical professional. / Response 1 miscalculates the break-even (140 days vs. ~179), a factual error. Response 2 is accurate, covers all points, and is more useful. Both follow constraints and are clear.
tieMaths word problemReasoning · 9.7 v 9.7Both correct, identical logic and answer, both clear and well structured. Response 2 slightly more formatted but not meaningfully better.
A tank holds 2,400 litres. Pump A fills at 40 L/min. Pump B drains at 25 L/min. A runs from 09:00. B accidentally switches on at 09:20. At 10:00 B is switched off. At what time is the tank full? Show your working.
The tank has a capacity of \(2400\) L. Pump A fills at a rate of \(40\) L/min and starts at \(09:00\). Pump B drains at a rate of \(25\) L/min; it is on only from \(09:20\) to \(10:00\). **Interval \(09:00\)–\(09:20\)** (only A running) Duration \(= 20\) min. Volume added \(= 20 \times 40 = 800\) L. Volume in tank at \(09:20\): \(800\) L. **Interval \(09:20\)–\(10:00\)** (A and B both running) Net rate \(= 40 - 25 = 15\) L/min. Duration \(= 40\) min. Volume added \(= 40 \times 15 = 600\) L. Volume in tank at \(10:00\): \(800 + 600 = 1400\) L. **After \(10:00\)** (only A running) Remaining volume \(= 2400 - 1400 = 1000\) L. Time required \(= 1000 / 40 = 25\) min. Thus the tank is full at \(10:25\). Final answer: \(10:25\)
To find the time when the tank is full, we can break the process down into three distinct time intervals based on which pumps are running. ### **Interval 1: 09:00 to 09:20 (20 minutes)** * **Active Pump:** Only Pump A (fills at 40 L/min). * **Water added:** $$\text{Time} \times \text{Rate} = 20 \text{ mins} \times 40 \text{ L/min} = 800 \text{ litres}$$ * **Volume in tank at 09:20:** **800 litres** --- ### **Interval 2: 09:20 to 10:00 (40 minutes)** * **Active Pumps:** Pump A (fills at 40 L/min) and Pump B (drains at 25 L/min). * **Net filling rate:** $$40 \text{ L/min} - 25 \text{ L/min} = 15 \text{ L/min}$$ * **Water added:** $$40 \text{ mins} \times 15 \text{ L/min} = 600 \text{ litres}$$ * **Volume in tank at 10:00:** $$800 \text{ L} + 600 \text{ L} = 1,400 \text{ litres}$$ --- ### **Interval 3: From 10:00 onwards** * **Active Pump:** Only Pump A (fills at 40 L/min). * **Remaining volume to fill:** $$\text{Total Capacity} - \text{Current Volume} = 2,400 \text{ L} - 1,400 \text{ L} = 1,000 \text{ litres}$$ * **Time needed to fill the remaining volume:** $$\frac{1,000 \text{ litres}}{40 \text{ L/min}} = 25 \text{ minutes}$$ --- ### **Conclusion** Adding 25 minutes to 10:00: The tank is full at **10:25**.
anthropic/claude-sonnet-5: tie — Both correct, identical logic and answer, both clear and well structured. Response 2 slightly more formatted but not meaningfully better. / Both responses correctly compute the same result (10:25) with identical clear step-by-step working. Equally accurate, clear, and useful; no constraints violated. | openai/gpt-5.6-terra: tie — Both responses correctly calculate 800 L by 09:20, then 600 L net during the 40-minute overlap, leaving 1,400 L at 10:00. Each correctly finds 1,000 L remaining, requiring 25 minutes at 40 L/min, so the tank is full at 10:25. Both show clear working; Response 1 is slightly more concise. / Both responses correctly calculate 800 L by 09:20, then 600 L net added by 10:00, leaving 1,000 L. At 40 L/min this takes 25 minutes, so the tank is full at 10:25. Both show clear working and meet all task requirements. | deepseek/deepseek-v4-pro: tie — Both responses are factually correct, show all working steps, and are clear and concise. No constraints violated. They are equally useful. / Both responses correctly calculate the fill time as 10:25, show clear working, and follow all instructions. No factual errors or constraint breaches. They are equally useful, clear, and concise.
tieCatch the flawReasoning · 9.2 v 9Both correctly identify correlation/causation and linearity flaws within word limits, clear and concise. Response 2 slightly more detailed but…
Find the flaw in this argument and explain it in under 100 words: 'Our churn dropped from 6% to 3% after we introduced the loyalty programme in March, so the programme cut churn in half. We should double the loyalty budget to cut churn to 1.5%.'
**Flaw: false causation + false linearity.** The drop after March shows only correlation, not that the programme caused it (other factors may explain the change). Even if it did, halving once does not imply doubling spend halves churn again—effects are rarely linear (diminishing returns, saturation, or a natural floor apply).
This argument contains two major flaws: 1. **Correlation vs. Causation:** It assumes the loyalty program caused the drop, ignoring external factors like seasonality, product updates, or competitor behavior. 2. **Diminishing Returns:** It assumes a linear relationship between budget and churn reduction. Doubling the budget will not automatically halve churn again. The remaining 3% of churning customers are likely much harder to retain, meaning further spending will yield diminishing returns, not a guaranteed 1.5% rate.
anthropic/claude-sonnet-5: tie — Both correctly identify correlation/causation and linearity flaws within word limits, clear and concise. Response 2 slightly more detailed but both are accurate and well within 100 words. / Both correctly identify correlation/causation and false linearity/diminishing returns flaws, are concise, under 100 words, and clear. Nearly identical quality; Response 1 slightly more detailed, Response 2 slightly more concise. | openai/gpt-5.6-terra: A — Both correctly identify unsupported causation and the unjustified assumption that doubling budget will proportionally reduce churn. Both are under 100 words. Response 1 is more concise and avoids Response 2’s speculative claim that remaining churners are likely harder to retain. / Both correctly identify post hoc causation and the unsupported assumption of linear budget effects, and both stay under 100 words. Response 2 is more concise, directly framed around the core inferential flaws, and is especially suitable for a busy professional. Response 1 is accurate but slightly more verbose. | deepseek/deepseek-v4-pro: tie (flipped on swap) — Both responses correctly identify correlation vs. causation and diminishing returns. They are accurate, concise, and useful for a non-technical reader, with no instruction violations. / Response 1 correctly identifies correlation/causation and diminishing returns with concrete examples, making it more useful for a non-technical reader. Response 2 is accurate but less detailed. Both under 100 words.
Grok 4.5Messy to JSONExtraction · 8.7 v 5Both extract accurate data; Response 2 wraps JSON in markdown code fences, violating 'return only valid JSON' instruction. Response 1 complies…
Extract to JSON with keys name, company, phone, email, intent: 'hiya - jules moreau here from Atelier Nord (the lighting people). best number is 07700 900 457, or jules@ateliernord.fr. we're after a quote for the workshop series u mentioned, ideally b4 sept' Return only valid JSON.
{ "name": "jules moreau", "company": "Atelier Nord", "phone": "07700 900 457", "email": "jules@ateliernord.fr", "intent": "quote for the workshop series ideally before sept" }
```json { "name": "Jules Moreau", "company": "Atelier Nord", "phone": "07700 900 457", "email": "jules@ateliernord.fr", "intent": "Requesting a quote for the workshop series, ideally before September" } ```
anthropic/claude-sonnet-5: A — Both extract accurate data; Response 2 wraps JSON in markdown code fences, violating 'return only valid JSON' instruction. Response 1 complies fully with pure JSON output, making it more instruction-compliant despite lowercase name. / Response 1 wraps JSON in markdown code fences, violating 'return only valid JSON' instruction. Response 2 complies exactly, though uses lowercase name. Both extract data correctly; Response 2 better follows format constraint. | openai/gpt-5.6-terra: A — Response 1 is valid JSON, includes all required keys, and accurately extracts the contact details and intent. Response 2’s extracted content is equally accurate and clearer in capitalization, but the Markdown code fence means it does not return only valid JSON as required. / Both extract all fields accurately. Response 1 wraps the JSON in Markdown code fences, so it does not comply with the instruction to return only valid JSON. Response 2 is valid JSON and accurately captures the contact details and quote request deadline. | deepseek/deepseek-v4-pro: A — Response 1 extracts correctly but uses lowercase. Response 2 capitalizes name, expands date, more useful. Response 2 violates return-only-JSON by wrapping in markdown; cap applies. Scores reflect quality ignoring that breach. / Response 1 capitalises name, expands abbreviations (b4→before, sept→September) making intent clearer for a busy professional. Both extract all fields correctly. Response 2 retains lowercased name and 'sept', slightly less clear. Violation: Response 1 wraps JSON in a code block, breaking 'Return only valid JSON' constraint.
tieTable from proseExtraction · 9.7 v 9.5Both tables are accurate and correctly formatted with matching percentages. Response 2 adds a brief intro and alignment formatting, slightly more…
Turn this into a markdown table with columns Product, Q1, Q2, Change: 'The Alpha line did 14,200 units in Q1 and 15,800 in Q2. Beta slipped from 9,100 to 7,400. The new Gamma launched mid-Q2 with 3,300 units (no Q1 sales). Delta held flat at 5,000 both quarters.' Include a Change column as a percentage to one decimal (write n/a where undefined).
| Product | Q1 | Q2 | Change | |---------|------|------|--------| | Alpha | 14,200 | 15,800 | 11.3% | | Beta | 9,100 | 7,400 | -18.7% | | Gamma | 0 | 3,300 | n/a | | Delta | 5,000 | 5,000 | 0.0% |
Here is the data formatted as a markdown table: | Product | Q1 | Q2 | Change | | :--- | ---: | ---: | ---: | | Alpha | 14,200 | 15,800 | +11.3% | | Beta | 9,100 | 7,400 | -18.7% | | Gamma | 0 | 3,300 | n/a | | Delta | 5,000 | 5,000 | 0.0% |
anthropic/claude-sonnet-5: tie — Both tables are accurate and correctly formatted with matching percentages. Response 2 adds a brief intro and alignment formatting, slightly more polished but unnecessary. Both meet constraints equally well. / Both correctly compute percentages and format tables identically in substance; Response 1 adds a '+' sign for clarity, a minor stylistic difference. Both meet all constraints and are equally clear and useful. | openai/gpt-5.6-terra: tie — Both responses accurately extract all quarterly values and calculate changes correctly to one decimal place. Both use n/a for Gamma’s undefined percentage change and provide the required markdown table columns. Response 2 adds an unnecessary but harmless introductory sentence; otherwise both are clear and useful. / Both accurately extract all quarterly figures, calculate percentage changes correctly to one decimal, use n/a for undefined Gamma change, and provide the requested markdown table columns. Response 1’s positive sign is acceptable; its brief introduction does not reduce usefulness. | deepseek/deepseek-v4-pro: tie (flipped on swap) — Response 1 is perfectly accurate, directly follows instructions without extraneous text, and is maximally concise. Response 2 adds a sentence and a positive sign (acceptable but not required), slightly less concise. / Both tables correctly present the data with accurate percentages and n/a. Response 2 is slightly more concise, but both fully meet requirements. No constraint violations.
tieContact detailsExtraction · 9.7 v 9.2Both responses are nearly identical, correctly listing all four people with role and org, using 'unknown' where appropriate. Minor wording…
List every person mentioned below with their role and organisation, one line each in the format Name - Role - Org. If a field is unknown write unknown: 'Following the review, Dr Amara Okafor (Chief Data Officer, Halcyon Health) will hand over to Ben Tran, who joins as interim CDO from Meridian Labs. Legal sign-off sits with Priya Shah at Halcyon; the external audit remains with T. Lindqvist.'
Dr Amara Okafor - Chief Data Officer - Halcyon Health Ben Tran - interim CDO - Meridian Labs Priya Shah - Legal sign-off - Halcyon T. Lindqvist - external audit - unknown
Dr Amara Okafor - Chief Data Officer - Halcyon Health Ben Tran - interim CDO - Meridian Labs Priya Shah - Legal sign-off - Halcyon T. Lindqvist - External Auditor - unknown
anthropic/claude-sonnet-5: tie — Both responses are nearly identical, correctly listing all four people with role and org, using 'unknown' where appropriate. Minor wording difference for Lindqvist's role doesn't affect accuracy or usefulness. / Both responses are nearly identical, correctly listing all four people with roles and orgs in the required format. Minor wording difference (External Auditor vs external audit) is negligible. | openai/gpt-5.6-terra: A — Response 1 accurately preserves each stated role and organisation, including unknown for Lindqvist’s organisation. Response 2 is otherwise correct but changes the stated role “external audit” to “External Auditor,” an unsupported interpretation. Both follow the required one-line format. / Both list all four people in the required three-field, one-line format. Response 2 most faithfully preserves T. Lindqvist’s stated responsibility as “external audit.” Response 1’s “External Auditor” is a reasonable inference but changes the wording from the source. | deepseek/deepseek-v4-pro: tie (flipped on swap) — Both responses correctly list all four persons with accurate roles and organizations, follow the prescribed format, and write 'unknown' for T. Lindqvist's org. They are equally accurate, compliant, useful, and clear. / Response 2 uses the exact phrase 'external audit' from the text, while Response 1 changes it to 'External Auditor', a minor inaccuracy. Both follow the format and are clear.
tieKey factsExtraction · 9.7 v 9.7Both responses correctly extract 5 concise bullets under 15 words, accurate and clear. Nearly identical quality, differing only in bullet style…
Extract exactly 5 key facts as bullets (each under 15 words) from: 'The council approved the harbour redevelopment on 12 June by 7 votes to 2. Phase one, costing £4.6m, begins in October and includes the fish market roof. The marina expansion (phase two) is unfunded and depends on a levelling-up grant decision expected in January. Local traders get first refusal on the 14 new retail units. Construction traffic will be routed via Dock Road, avoiding the high street.'
- Council approved harbour redevelopment on 12 June by 7-2. - Phase one costs £4.6m, starts October, includes fish market roof. - Marina expansion unfunded, awaits January levelling-up grant decision. - Local traders get first refusal on 14 new retail units. - Construction traffic routes via Dock Road, avoids high street.
* The council approved the harbour redevelopment on 12 June by 7-2 votes. * Phase one costs £4.6m, begins in October, and includes the fish market roof. * Phase two marina expansion is unfunded, pending a January levelling-up grant decision. * Local traders have first refusal on the 14 new retail units. * Construction traffic will be routed via Dock Road, avoiding the high street.
anthropic/claude-sonnet-5: tie — Both responses correctly extract 5 concise bullets under 15 words, accurate and clear. Nearly identical quality, differing only in bullet style and minor phrasing. / Both extract exactly 5 accurate bullets under 15 words each, following instructions well. Response 2 is slightly more concise, but both are equally correct and useful. | openai/gpt-5.6-terra: tie — Both provide exactly five bullet points, each under 15 words, and accurately capture all key facts. Response 2 is marginally more explicit about phase two, while Response 1 is slightly more concise; neither has a meaningful advantage. / Both provide exactly five bullets, each under 15 words, and accurately capture all five key facts. Response 1 is slightly fuller; Response 2 is slightly more concise. Both are clear and useful. | deepseek/deepseek-v4-pro: tie — Both responses extract 5 key facts accurately, with bullets under 15 words each, no constraint violations, equally clear and useful for a professional. / Both responses have exactly 5 factually accurate bullets, each under 15 words. They meet all instructions, are clear and concise, and equally useful for a busy professional.
tieSchema complianceExtraction · 9.8 v 9.8Both responses produce identical, correct JSON matching the schema exactly, differing only in formatting whitespace, which is inconsequential.
Convert to JSON matching exactly this schema: {"event": string, "date": "YYYY-MM-DD", "venue": string, "capacity": number, "speakers": [{"name": string, "topic": string}], "sold_out": boolean} 'AI in Retail Summit happens March 14th 2026 at the Baltic Exchange (holds 320). Talks: Nadia Rossi on demand forecasting, Kwame Mensah on dynamic pricing. Tickets still available.' Return only the JSON.
{ "event": "AI in Retail Summit", "date": "2026-03-14", "venue": "Baltic Exchange", "capacity": 320, "speakers": [ {"name": "Nadia Rossi", "topic": "demand forecasting"}, {"name": "Kwame Mensah", "topic": "dynamic pricing"} ], "sold_out": false }
{ "event": "AI in Retail Summit", "date": "2026-03-14", "venue": "Baltic Exchange", "capacity": 320, "speakers": [ { "name": "Nadia Rossi", "topic": "demand forecasting" }, { "name": "Kwame Mensah", "topic": "dynamic pricing" } ], "sold_out": false }
anthropic/claude-sonnet-5: tie — Both responses produce identical, correct JSON matching the schema exactly, differing only in formatting whitespace, which is inconsequential. / Both responses are correct, valid JSON matching the schema exactly, differing only in formatting style. Both are equally clear and accurate. | openai/gpt-5.6-terra: tie — Both responses are valid JSON and exactly match the required schema. All extracted details are accurate: event, ISO date, venue, numeric capacity, speaker names/topics, and sold_out false from tickets being available. Formatting differences do not affect correctness or usability. / Both responses are valid JSON, match the required schema exactly, accurately extract all event details, and correctly infer sold_out as false from tickets being available. Both comply with the instruction to return only JSON. | deepseek/deepseek-v4-pro: tie — Both responses are identical in content and fully compliant with the schema and instructions. The minor formatting difference (inline vs. multi-line) does not impact correctness or clarity. / Both responses are valid JSON that match the schema exactly and contain all correct data. No constraints were set or broken. They are identical in content, making them equally correct, useful, and clear.
Gemini 3.5 FlashFormat constraintsInstruction-following · 8 v 8.5Response 1 follows all constraints: 3 sentences, correct starting words, under 20 words each. Response 2 violates the 20-word limit in sentence 1…
Describe how photosynthesis works in exactly 3 sentences. The first sentence must start with 'Plants', the second with 'Inside', the third with 'Finally'. No sentence may exceed 20 words.
Plants convert sunlight water and carbon dioxide into glucose plus oxygen. Inside chloroplasts light energy splits water molecules releasing oxygen gas. Finally carbon dioxide combines with hydrogen forming sugars for plant energy.
Plants absorb sunlight, water, and carbon dioxide to initiate the process of making their own food. Inside chloroplasts, light energy converts these raw ingredients into chemical energy, specifically glucose. Finally, oxygen is released into the air as a byproduct while the plant uses glucose for growth.
anthropic/claude-sonnet-5: tie (flipped on swap) — Response 1 follows all constraints: 3 sentences, correct starting words, under 20 words each. Response 2 violates the 20-word limit in sentence 1 (16 words, ok actually) - checking: sentence1=16 words ok, sentence2=17 words ok, sentence3=20 words ok. Actually both seem within limit; but response 2 sentences longer, less concise. Response 1 more accurate and concise. / Both meet sentence structure and word limits, but R1 is more accurate; R2's third sentence describes photosynthesis somewhat inaccurately (carbon dioxide combining with hydrogen) and lacks punctuation clarity. | openai/gpt-5.6-terra: B — Both meet the three-sentence, required-opening, and under-20-word constraints. Response 2 is more accurate and accessible, clearly connecting glucose to growth and oxygen to byproduct release. Response 1 is concise and broadly correct but oversimplifies sugar formation as carbon dioxide combining directly with hydrogen. / Both meet all sentence, opening-word, and word-limit constraints. Response 1 is clearer and more complete for a non-technical reader, accurately linking inputs, chloroplasts, glucose, and oxygen. Response 2 is concise and broadly accurate but is less polished and omits punctuation between listed inputs. | deepseek/deepseek-v4-pro: B — Both follow constraints perfectly. Response 2 is more polished and clearer for a non-technical professional, using proper punctuation and phrasing like 'chemical energy'. Response 1's first sentence lacks commas, slightly reducing clarity. Accuracy is comparable, but 2 is marginally better. / Both responses meet constraints. Response 1 is more articulate and factual, while Response 2 is slightly less clear and oversimplified.
tieWord limitInstruction-following · 9.3 v 8.5Both responses are exactly 50 words, accurate, clear, and cover key distinctions (exclusive possession, proprietary interest vs personal…
Explain the difference between a lease and a licence for UK commercial property in exactly 50 words. Count carefully - exactly 50.
A lease grants exclusive possession of UK commercial property for a fixed term, creating a legal interest in land enforceable against others. A licence only allows use without exclusive possession, being a personal permission not creating property rights. Leases offer greater security and protections than licences under UK property law.
A commercial lease grants exclusive possession of property for a fixed term, creating a proprietary interest that binds future owners. Conversely, a licence merely provides personal permission to occupy the space, offering no exclusive possession or security of tenure, and can be revoked easily by the landlord at any time.
anthropic/claude-sonnet-5: tie — Both responses are exactly 50 words, accurate, clear, and cover key distinctions (exclusive possession, proprietary interest vs personal permission). Equally useful and concise; no meaningful difference in quality. / Both responses are exactly 50 words, factually accurate, clear, and cover the key lease/licence distinctions. Both equally useful and concise, so this is effectively a tie. | openai/gpt-5.6-terra: A — Both responses contain exactly 50 words. Response 1 accurately captures the core distinction and is clear, though leases need not always be fixed-term. Response 2 is clear but overstates that licences can be revoked by landlords “at any time”; contractual licence terms may limit revocation. / Both are exactly 50 words and accurately state the core distinction. Response 2 is slightly stronger: it avoids Response 1's overbroad suggestion that a licence can always be revoked at any time, which depends on its terms. Both are clear and useful. | deepseek/deepseek-v4-pro: tie (flipped on swap) — Response 1 is fully correct and balanced. Response 2 inaccurately states a licence 'can be revoked easily at any time,' ignoring fixed-term contractual licences that offer some stability. / Both responses correctly explain the difference: lease grants exclusive possession and proprietary interest; licence is personal permission without security. Each is exactly 50 words, clear, and useful. No violations.
tieForbidden wordsInstruction-following · 8.2 v 8.5Both responses stay within word count, avoid banned words and variants, and explain bitterness accurately. Response 2 is slightly more detailed…
Explain what makes coffee taste bitter, in 80-120 words, WITHOUT using any of these words: bitter, bean, roast, caffeine, acid. Do not use hyphenated or partial variants of them either.
The sharp and astringent quality in coffee arises mainly from over-extraction during brewing. When ground particles stay in contact with water longer than ideal, too many soluble substances dissolve, producing that unwanted edge. Elevated water temperatures speed up the release of these elements. Additionally, the heating treatment applied to the seeds creates complex molecules responsible for the intensity. A finer grind size increases exposure and worsens the issue. To avoid this, monitor brew time closely, use moderate heat, and select a coarser texture for balance. Freshness of the grounds also plays a role in overall smoothness.
When hot water extracts flavors from ground coffee seeds, timing is everything. If the water stays in contact with the grounds for too long, or if the water is too hot, it pulls out heavy organic compounds. These include substances like phenylindanes, which develop when the seeds are cooked at high temperatures. While initial extraction yields pleasant, sweet, and fruity notes, continued brewing releases these harsher, sharper chemicals. Additionally, grinding the seeds too finely increases the surface area, leading to over-extraction. This results in a harsh, dry sensation on the tongue, masking the smoother flavors and leaving a strong, sharp aftertaste.
anthropic/claude-sonnet-5: tie (flipped on swap) — Both responses stay within word count, avoid banned words and variants, and explain bitterness accurately. Response 2 is slightly more detailed and engaging, but both are clear, concise, and equally useful. / Both avoid banned words and meet word count. Response 1 offers more specific, accurate detail (phenylindanes, grind size) while Response 2 is slightly more generic but still clear and useful. | openai/gpt-5.6-terra: B — Both meet the word-count and forbidden-term constraints. Response 1 is clear and useful but overstates over-extraction as the main cause and is less specific about the underlying compounds. Response 2 more accurately explains extraction, heat-created phenylindanes, grind size, and sensory effects while remaining accessible. / Both meet the 80–120-word range and avoid all prohibited terms and variants. Response 1 gives a more accurate, specific explanation of over-extraction, heat, fine grinding, and phenylindanes. Response 2 is clear and practical but less precise, relying on vague references to “elements” and “complex molecules.” | deepseek/deepseek-v4-pro: tie (flipped on swap) — Response 1 offers practical tips (brew time, heat, grind) directly useful to a busy professional. Both are accurate and follow constraints, but Response 1 is slightly more concise and actionable. / Both correctly attribute bitterness to over-extraction, avoid forbidden terms, and stay within word limits. Response 1 cites phenylindanes; Response 2 emphasizes practical tips. Equally clear and useful.
Gemini 3.5 FlashMulti-constraintInstruction-following · 5 v 9.2Response1 is only 85 words, violating the 90-110 word count; Response2 meets all constraints (word count, 3 paragraphs, single-question paragraph…
Write a LinkedIn post about hiring your first apprentice. Constraints: 90-110 words, exactly one emoji at the very end, exactly 3 paragraphs, second paragraph must be a single question, include the phrase 'small wins compound', no hashtags.
Hiring my first apprentice marked a turning point in building our team culture and future talent pipeline. The process taught me patience, clearer communication, and the joy of watching someone grow from novice to valued contributor through daily guidance and real project ownership. What surprised you most about mentoring your initial apprentice hire? Over months those early lessons stuck and small wins compound into lasting confidence for everyone involved. This experience reaffirmed why investing in people creates stronger organizations long term while giving back meaningfully. 🌟
Today marks a truly significant milestone for our growing team as we officially welcome our very first apprentice. Bringing on new talent is not just about getting extra help; it is about actively investing in the future of our industry. We are incredibly excited to mentor, guide, and learn from this fresh perspective. Do you remember when someone first took a chance on you at the very beginning of your professional journey? By focusing on daily growth, we know that small wins compound into incredible lasting success. We are committed to building a supportive environment where practical learning is the top priority. Let the journey begin. 🚀
anthropic/claude-sonnet-5: B — Response1 is only 85 words, violating the 90-110 word count; Response2 meets all constraints (word count, 3 paragraphs, single-question paragraph, phrase included, one emoji, no hashtags) and reads naturally, making it more accurate and useful. / Both meet structure, phrase, single question, one emoji, no hashtags. R1 is within 90-110 words (104); R2 is only 85 words, violating the word count constraint. R1 is otherwise well-crafted and professional. | openai/gpt-5.6-terra: B — Response 1 is well-written and meets the paragraph, question, phrase, emoji, and hashtag requirements, but has only 85 words. Response 2 meets every explicit constraint: 106 words, three paragraphs, a single-question second paragraph, required phrase, no hashtags, and one final emoji. It is clear, professional, and suitable for LinkedIn. / Response 1 meets all constraints: 105 words, three paragraphs, a single-question second paragraph, required phrase, no hashtags, and one final emoji. It is polished and useful. Response 2 has strong content and structure but contains only 85 words, violating the required 90–110-word range; its opening of paragraph three is also slightly awkward. | deepseek/deepseek-v4-pro: B — Response 1 is more personal, concise, and resonant for a busy professional; includes required phrase and structure. It violates word count (85 words) which will be capped. Response 2 complies fully but is slightly generic. Content-wise, 1 is stronger. / Response 1 perfectly meets all constraints with engaging, clear content. Response 2 has good content but fails word count (85 words). Scored ignoring that breach as instructed, but violation reported.
tieExact structureInstruction-following · 9 v 9.2Both follow structure with title, three sections, 4 items each, 3-8 words per item. Response 1 has no blank lines between sections while Response…
Produce a checklist for closing a small retail shop at night with EXACTLY this structure: a title line, then 3 sections headed 'Front of house', 'Back office', 'Security', each containing exactly 4 numbered items, each item 3-8 words. Nothing else before or after.
Small Retail Shop Night Closing Checklist Front of house 1. Turn off all display lights 2. Lock the main entrance 3. Tidy up sales floor area 4. Empty customer trash bins Back office 1. Count and secure all cash 2. Shut down computer systems 3. File daily sales paperwork 4. Lock all office cabinets Security 1. Activate the alarm system 2. Check every window lock 3. Verify cameras are recording 4. Secure the rear exit door
Nightly Retail Closing Checklist Front of house 1. Clear and wipe down all counters. 2. Restock the main display shelves. 3. Sweep and mop the sales floor. 4. Turn off all display lights. Back office 1. Count the cash register drawer. 2. Balance the daily sales report. 3. File the daily transaction receipts. 4. Lock the safe securely. Security 1. Check all back exit doors. 2. Turn on the night lights. 3. Activate the main burglar alarm. 4. Lock the front entrance door.
anthropic/claude-sonnet-5: tie (flipped on swap) — Both follow structure with title, three sections, 4 items each, 3-8 words per item. Response 1 has no blank lines between sections while Response 2 has blank lines, but task didn't forbid this. Both are equally correct and useful. / Response 1 follows structure with title line separated, proper spacing, all items 3-8 words. Response 2 lacks blank lines between title/sections and among sections, deviating from clean structure, though content is fine. Both meet word counts, but Response 2's formatting is more cramped. | openai/gpt-5.6-terra: tie (flipped on swap) — Both responses use a title followed by the three required headings, each with exactly four numbered items. Every item contains 3–8 words, and both provide practical, accurate retail-closing tasks. Response 2 includes blank spacing but no extraneous content. / Both exactly follow the required structure and item word counts. Response 2 is slightly more practical and concise, covering cash security, systems, cameras, windows, and exits. Response 1 is also strong but includes less comprehensive security checks. | deepseek/deepseek-v4-pro: A — Both follow the structure exactly. Response 1 has no typos and slightly more thorough security items (cameras, window locks). Response 2 has a minor spelling error ('burglar') and omits camera/window checks, making it marginally less useful. / Both follow the structure perfectly. Response 1 has a typo ('burglar' alarm) reducing factual accuracy, while Response 2 is flawless and equally useful.
tieQuick answer (short)Speed & cost · 10 v 10Both responses are correct, concise, one-sentence answers stating VAT meaning and 20% UK rate, essentially equivalent in quality.
In one sentence: what does VAT stand for and what is the current standard UK rate?
VAT stands for Value Added Tax, and the current standard UK rate is 20%.
VAT stands for Value Added Tax, and its current standard rate in the UK is 20%.
anthropic/claude-sonnet-5: tie — Both responses are correct, concise, one-sentence answers stating VAT meaning and 20% UK rate, essentially equivalent in quality. / Both responses correctly and concisely state VAT stands for Value Added Tax with a 20% standard UK rate, in a single sentence as required. Virtually identical in quality. | openai/gpt-5.6-terra: tie — Both responses are factually correct, use one sentence, fully answer both parts, and are clear and concise for a non-technical reader. / Both responses are factually correct, use one sentence, directly answer both parts, and are equally clear and concise for a non-technical professional. | deepseek/deepseek-v4-pro: tie — Both responses are factually correct, one sentence, and clear. No violations. / Both responses are one sentence, factually correct, and directly answer the question with the required information.
tieInstant classificationSpeed & cost · 9.5 v 9.3Both correctly classify all items. Response 1 uses clean format matching original text without quotes, slightly cleaner presentation than Response…
Classify each as SPAM or NOT SPAM, output only two-column lines: 'You've won a £500 voucher click here' / 'Your Tesco order 8812 is ready for collection' / 'URGENT your account will be closed verify now' / 'Minutes from Thursday's board meeting attached'.
You've won a £500 voucher click here / SPAM Your Tesco order 8812 is ready for collection / NOT SPAM URGENT your account will be closed verify now / SPAM Minutes from Thursday's board meeting attached / NOT SPAM
'You've won a £500 voucher click here' | SPAM 'Your Tesco order 8812 is ready for collection' | NOT SPAM 'URGENT your account will be closed verify now' | SPAM 'Minutes from Thursday's board meeting attached' | NOT SPAM
anthropic/claude-sonnet-5: tie (flipped on swap) — Both correctly classify all items. Response 1 uses clean format matching original text without quotes, slightly cleaner presentation than Response 2 which retains quotes and uses pipe instead of slash separator. / Both correctly classify all four items identically; only formatting differs slightly (pipe vs slash), neither strictly matching original slash format perfectly but Response 2 uses the same delimiter as task, making it marginally more compliant, though both are functionally equivalent. | openai/gpt-5.6-terra: tie — Both correctly classify all four messages and provide only two-column classification lines. Response 2 is slightly more explicitly tabular, but neither violates a required delimiter or format. / Both correctly classify the promotional and urgent verification messages as SPAM and the legitimate order and meeting messages as NOT SPAM. Each provides only four two-column classification lines; differing separators do not violate a specified format. | deepseek/deepseek-v4-pro: tie — Both responses correctly classify all four items, output only two-column lines as instructed, and are equally clear and concise. No constraints are violated. / Both correctly classify all phrases, output only two-column lines with no extra text, and are clear and concise for a non-technical reader.
tieShort translationSpeed & cost · 9.1 v 7.9Both accurate, informal translations. Response 1 is concise and direct. Response 2 offers extra alternatives, which is useful but less concise and…
Translate to French, keeping the informal tone: 'Hey, we're running 10 minutes late - grab us a table if you get there first?'
**Salut, on a 10 minutes de retard – prends-nous une table si t’arrives avant ?**
Here is a very natural, informal translation (perfect for a text message to a friend): **"Hey, on a 10 minutes de retard — tu nous prends une table si t'arrives en premier ?"** *Alternative (slightly more slangy/casual):* **"Salut, on a 10 min de retard. Tu nous chopes une table si t'arrives avant ?"**
anthropic/claude-sonnet-5: tie (flipped on swap) — Both accurate, informal translations. Response 1 is concise and direct. Response 2 offers extra alternatives, which is useful but less concise and slightly redundant for a simple task. / Both are accurate, informal French translations. Response 1 offers a primary translation plus a casual alternative, giving more value, while Response 2 is concise and correct but slightly less rich. Both follow instructions well. | openai/gpt-5.6-terra: A — Response 1 is concise and informal, but “prends-nous une table” is a somewhat literal, less idiomatic rendering; “garde-nous une table” would be more natural. Response 2 provides usable French options, especially “chopes une table,” but adds unnecessary English commentary and its main wording (“arrives en premier”) is less natural. / Response 2 is concise, natural French and preserves the informal tone and meaning. Response 1 offers workable translations, but adds unnecessary commentary and alternatives; “si t’arrives en premier” is slightly less idiomatic in this context than Response 2’s phrasing. | deepseek/deepseek-v4-pro: tie (flipped on swap) — Response 1 is fully correct, uses natural informal French ('Salut', 'prends-nous'), and is concise. Response 2 retains English 'Hey' (less accurate), adds unnecessary commentary, and is less efficient for a busy professional. / Both translations are accurate, use informal French ('on a', 't’arrives'), and respect the original meaning. Response 1 provides helpful alternatives, while Response 2 is a clear single option. No constraints were violated.
Gemini 3.5 FlashMedium generationSpeed & cost · 6 v 8.7Both meet requirements; Response 2 adds vivid named details (Rosewood, Arthur, Margaret, Sarah, parkin) making it more engaging and specific…
Write a 150-word 'About us' section for a family-run garden centre in Yorkshire founded in 1987, now run by the founders' daughter, known for rare perennials and a small cafe.
**About Us** Nestled in the heart of Yorkshire, our family-run garden centre has been a cherished local haven since 1987. Founded with passion and green-fingered dedication by our parents, it began as a modest plot nurturing plants and community spirit alike. Today, their daughter proudly continues the legacy, bringing fresh energy while honouring the values that have defined us for nearly four decades. We specialise in rare and unusual perennials, carefully sourced and grown to delight discerning gardeners seeking something beyond the ordinary. From elusive heritage varieties to striking contemporary finds, our collections inspire borders that truly stand out through every season. Complementing the greenery is our cosy on-site café, where visitors can relax with homemade cakes, fresh coffee, and light lunches amid the scent of flowers and the gentle Yorkshire countryside. Whether you’re a seasoned horticulturist or simply love a peaceful browse, we welcome you to explore, discover, and feel at home. Come and share in our family’s enduring love of plants.
Nestled in the heart of the Yorkshire countryside, Rosewood Gardens has been a haven for plant lovers since 1987. Founded by Arthur and Margaret, our family-run garden centre is now proudly managed by their daughter, Sarah, who carries on their lifelong passion for all things green. We are renowned across the region for our exceptional collection of rare perennials, carefully nurtured to thrive in our northern climate. Whether you are a seasoned horticulturist or a budding gardener, our knowledgeable team is always on hand with friendly, expert advice. After exploring our vibrant plant areas, take a moment to relax in our cozy, independent little café. Enjoy a freshly brewed coffee and a slice of homemade Yorkshire parkin, made using local ingredients. At Rosewood, we’re more than just a garden centre; we’re a community of passionate plant lovers. Come visit us today and let us help your own beautiful garden grow.
anthropic/claude-sonnet-5: B — Both meet requirements; Response 2 adds vivid named details (Rosewood, Arthur, Margaret, Sarah, parkin) making it more engaging and specific, while still concise and accurate. Response 1 is generic but well-written. Word counts close to 150 for both. / Both meet the brief well; Response 1 is close to 150 words (~155) with specific names adding warmth, while Response 2 is longer (~170 words), deviating more from the target. Both clear and useful; Response 1 slightly more concise and accurate to word count. | openai/gpt-5.6-terra: tie (flipped on swap) — Response 1 is warmer, more polished, and covers all requested details without unnecessary invented business names or founder identities. Response 2 is clear and useful but adds unsupported specific names and details. Both miss the explicit 150-word requirement: Response 1 has 160 words (excluding heading), and Response 2 has 147. / Response 1 is exactly 150 words and clearly covers the Yorkshire location, 1987 founding, daughter-led family ownership, rare perennials, and café. Response 2 is polished and relevant but is 161 words excluding its heading (163 including it), so it misses the explicit word-count requirement. | deepseek/deepseek-v4-pro: B — Response 1 exceeds 150 words (161), violating constraint. Both are accurate and useful, but Response 2 is exactly 150 words, includes specific names and local detail, making it slightly more engaging and perfectly compliant. / Response 1 is more personal, naming founders and daughter, and includes a specific local treat. Response 2 is well-written but slightly less vivid. Both are accurate and useful, but Response 1 is more engaging.
tieLong generationSpeed & cost · 6.2 v 6.2Both are accurate, thorough, and well-organized covering all required points. Response 1 explicitly meets the 400-word count (412 words) while…
Write a detailed 400-word beginner's guide to setting up a home compost bin: location, what to add, what to avoid, how long it takes, and one common mistake.
**Beginner's Guide to Setting Up a Home Compost Bin** Composting turns kitchen scraps and yard waste into nutrient-rich soil amendment, reducing trash and benefiting your garden. This simple guide covers the essentials for beginners. **Choosing the Right Location** Select a dry, shady spot with good drainage—ideally near a water source and your kitchen or garden for convenience, but not right against the house to avoid pests or odors. A level area of about 3x3 feet works for most bins. Ensure easy access year-round and some airflow; avoid low spots that flood or full sun that dries the pile too fast. Place it on bare soil if possible so worms and microbes can enter naturally. Urban dwellers can use a compact enclosed bin on a balcony or patio with a tray underneath. **What to Add** Aim for a balance of “greens” (nitrogen-rich, moist materials) and “browns” (carbon-rich, dry materials), roughly equal volumes or a 1:2–1:3 green-to-brown ratio by volume. Greens include fruit and vegetable scraps, coffee grounds and filters, tea bags, fresh grass clippings, and plant trimmings. Browns include dry leaves, shredded cardboard or newspaper (non-glossy), straw, small twigs, sawdust from untreated wood, and paper egg cartons. Chop larger items into smaller pieces to speed breakdown. Add materials in layers or mix them, and keep the pile as moist as a wrung-out sponge—water lightly if dry, or add browns if soggy. **What to Avoid** Never add meat, bones, fish, dairy, oils, or greasy foods—these attract pests and create odors. Skip pet waste (dog/cat feces), diseased plants, weeds gone to seed, coal ash, treated wood, glossy or colored paper/magazines, and synthetic materials. Citrus and onions in moderation are usually fine, but large amounts can slow microbes. When in doubt, leave it out. **How Long It Takes** With regular turning (every 1–2 weeks for aeration) and proper moisture/balance, you can have usable compost in 2–6 months. Hot composting (pile heats to 130–160°F) finishes faster (as little as 1–3 months); cooler passive piles take longer, up to a year. Finished compost is dark, crumbly, and earthy-smelling with no recognizable scraps. Sift if needed and return unfinished bits to the bin. **One Common Mistake** A frequent beginner error is adding too many greens without enough browns, creating a slimy, smelly anaerobic mess. Always bulk up with dry leaves or cardboard and turn the pile to introduce oxygen—this prevents problems and keeps the process efficient and odor-free. Start small, observe, and adjust. In time you’ll produce free fertilizer while cutting waste. Happy composting! *(Word count: 412)*
Transforming kitchen scraps and yard waste into nutrient-rich "black gold" is one of the most rewarding eco-friendly habits you can adopt. Starting a home compost bin is simple, requiring just a little space and a basic understanding of the composting recipe. Here is how to get started. ### 1. Choosing the Perfect Location Place your compost bin directly on bare soil rather than concrete or decking. This allows beneficial microbes, earthworms, and fungi to easily migrate from the ground up into your pile. Select a flat, well-drained spot that receives partial shade. Too much direct sun will dry the pile out, while constant shade can keep it too cold and soggy. Lastly, ensure it is easily accessible from your kitchen so discarding scraps never feels like a chore. ### 2. What to Add (The Recipe) Successful composting relies on a balanced diet of "greens" and "browns." For a healthy pile, aim for a ratio of roughly two to three parts browns to one part greens. * **Greens (Nitrogen-rich):** These provide moisture and nutrients. Include fruit and vegetable scraps, coffee grounds, tea leaves, and fresh grass clippings. * **Browns (Carbon-rich):** These create structure, allowing oxygen to flow. Include dry leaves, shredded cardboard, non-glossy paper, twigs, and straw. ### 3. What to Avoid To prevent pests and unpleasant odors, keep certain items out of your bin. Avoid meat, fish, dairy, bones, fats, and cooking oils, as these will attract rodents and raccoons. Do not add dog or cat waste, which can harbor harmful pathogens. Finally, skip diseased plants, weeds with mature seed heads, and chemically treated lawn clippings, as these can survive the composting process and contaminate your future garden. ### 4. How Long It Takes Composting is a lesson in patience. Depending on the weather, the size of your pile, and how often you tend to it, your compost will be ready in anywhere from **three months to a year**. You will know it is finished when the original ingredients are unrecognizable, and it looks like dark, crumbly soil with a pleasant, earthy smell. ### 5. One Common Mistake: Neglecting Aeration The most common beginner mistake is **forgetting to turn the pile**. Composting microbes require oxygen to break down organic matter. Without it, the pile becomes anaerobic, resulting in a compacted, slimy, and foul-smelling mess. To avoid this, use a pitchfork or compost aerator to turn and fluff your pile once every week or two, rotating the outer materials into the warm center. With just a little patience and the right balance, you’ll soon turn everyday waste into a free, nutrient-rich superfood for your garden.
anthropic/claude-sonnet-5: tie (flipped on swap) — Both are accurate, thorough, and well-organized covering all required points. Response 1 explicitly meets the 400-word count (412 words) while Response 2's length is unstated and appears longer, though not verified as violation. Response 1 is slightly more detailed and precise. / Both are accurate, complete, clear guides covering all required elements within word count. Response 1 is slightly more concise and better organized with headers; Response 2 says shade preferred but earlier guides often note partial shade—minor inconsistency (dry+shady is odd pairing). Both close, Response 1 slightly better polished. | openai/gpt-5.6-terra: tie — Both cover every requested topic accurately and use accessible beginner-friendly guidance. Response 1 is more detailed and practical, with stronger location, moisture, and troubleshooting advice. Response 2 is clear but somewhat less comprehensive. Response 1 explicitly states 412 words; Response 2 also does not meet the exact 400-word requirement. / Both accurately cover all requested topics in clear beginner-friendly language. Response 1 is well structured and practical; Response 2 adds useful moisture, bin-size, and apartment guidance. Response 2 has minor ratio ambiguity, while Response 1 is slightly less comprehensive. Neither is exactly 400 words (Response 2 states 412). | deepseek/deepseek-v4-pro: tie — Both guides are accurate and useful. Response 1 provides more detail (moisture, ratio, turning frequency, sifting) and better covers all points. Response 2 is clear but slightly less thorough. Both exceed 400 words. / Both responses are factually accurate, cover all required topics, and are well-structured for a busy professional. They are equally clear and useful. Word counts exceed 400, so both violate the constraint.
What they cost
- api price in per 1m$2OpenRouter API12 Aug 2026stale
- api price out per 1m$6OpenRouter API12 Aug 2026stale
- api price in per 1m$1.5OpenRouter API12 Aug 2026stale
- api price out per 1m$9OpenRouter API12 Aug 2026stale
Speed, measured
Median response time across all 30 tasks: 9.6s (Grok 4.5) vs 7.7s (Gemini 3.5 Flash).
Receipts — every citation, raw outputs
The full run — prompts, both raw outputs per task, judge reasoning, latencies, token counts and per-call cost — is downloadable as JSON. The task suite is versioned (suite-2026-07) and published on the methodology page.
- judge panelanthropic/claude-sonnet-5, openai/gpt-5.6-terra, deepseek/deepseek-v4-proour run (raw outputs)12 Aug 2026verified
- judge swap agreement0.467our run (raw outputs)12 Aug 2026verified
- judge swap kappa0.186our run (raw outputs)12 Aug 2026verified
- median latency ms a9578our run (raw outputs)12 Aug 2026verified
- median latency ms b7732our run (raw outputs)12 Aug 2026verified
- panel swap flip rate0.411our run (raw outputs)12 Aug 2026verified
- panel unanimous rate0.433our run (raw outputs)12 Aug 2026verified
- run cost a usd0.1079our run (raw outputs)12 Aug 2026verified
- run cost b usd0.4133our run (raw outputs)12 Aug 2026verified
- score a2our run (raw outputs)12 Aug 2026verified
- score b4our run (raw outputs)12 Aug 2026verified
- suite winswriting: a 1/b 1/tie 3 · coding: a 0/b 0/tie 5 · reasoning: a 0/b 0/tie 5 · extraction: a 1/b 0/tie 4 · instruction: a 0/b 2/tie 3 · speed cost: a 0/b 1/tie 4our run (raw outputs)12 Aug 2026verified
- tasks total30our run (raw outputs)12 Aug 2026verified
- ties24our run (raw outputs)12 Aug 2026verified
a narrow win on the tasks that separated them (6 of 30 tasks were decisive) — close enough that the loser is still worth a look.
“Grok 4.5 is a bigger deal than Fable 5 - YouTube”on Grok 4.5
“Gemini 3.6 Flash has been really good??”on Gemini 3.5 Flash