Battles / Best free
Gemini 3.5 Flash wins 4–3
a narrow win on the tasks that separated them (7 of 30 tasks were decisive) — close enough that the loser is still worth a look.
GLM 5.2 vs Gemini 3.5 Flash · Best freemodels in this battle: GLM 5.2 vs Gemini 3.5 Flash
Show the full workings
The evidence
Suite-by-suite
blue = GLM 5.2 wins · grey = ties · white = Gemini 3.5 Flash wins (5 tasks per suite)
Round-by-round — all 30 tasks
tieCold emailWriting · 8.7 v 8.3Both meet word limit, tone, and CTA requirements. Response 2 has slightly stronger personalization and clearer CTA (specific call ask), Response 1…
Write a cold email (max 120 words) from the owner of a 6-person UK accounting firm to a local brewery, offering a free first-month bookkeeping trial. Friendly but professional, one clear call to action, no buzzwords.
Subject: A free bookkeeping trial for [Brewery Name] Hi [Name], I’m [Your Name], owner of [Firm Name], a six-person accounting firm based here in [City/Town]. We love what you’re brewing and know that managing the books takes time away from the brewhouse. To show you how we work, I’d like to offer your first month of bookkeeping completely free. We handle daily transactions, payroll, and VAT, so you can focus on making great beer. If you’d like to take us up on this, just reply to this email and we can arrange a quick call. Cheers, [Your Name] Owner, [Firm Name]
Subject: Bookkeeping help for [Brewery Name] (from a local fan) Hi [Name], I’m [My Name], and I run a small, 6-person accounting team just down the road from you. We’re big fans of your beers—the [Beer Name] is a regular favourite in our office. We help local businesses keep their finances sorted. To help you spend less time on spreadsheets and more time brewing, we’d love to handle your bookkeeping for free for your first month. There’s no obligation or contract—just a simple way to see how much time we can save you. Are you free for a quick, 5-minute phone call next week to see if we're a good fit? Best regards, [My Name] [Firm Name]
anthropic/claude-sonnet-5: tie — Both meet word limit, tone, and CTA requirements. Response 2 has slightly stronger personalization and clearer CTA (specific call ask), Response 1 slightly more concise. Both well within 120 words, no buzzwords. / Both meet word limit, tone, and CTA requirements clearly. Response 1 adds a nice personal touch; Response 2 is slightly more informative about services. Both concise, professional, no buzzwords. Essentially equal quality. | openai/gpt-5.6-terra: tie (flipped on swap) — Both meet the length limit, offer a free first month, use a friendly professional tone, avoid obvious buzzwords, and include one clear CTA. Response 1 states ownership more explicitly and is slightly tighter. Response 2 feels more personalised and local, but “accounting team” is marginally less clearly an accounting firm. / Both are under 120 words, friendly, professional, and include one clear call to action. Neither explicitly states the firm is UK-based. Response 2 more clearly establishes that the sender is the owner and gives a simpler reply-based CTA. Response 1 is strong but is slightly less direct about ownership. | x-ai/grok-4.5: tie (flipped on swap) — Both under 120w, friendly-pro, one CTA, no buzzwords. R1 tighter, clearer VAT/UK fit, stronger usefulness. / Both meet all criteria: under 120 words, friendly-professional tone, clear free-trial offer, single CTA, no buzzwords. Equally correct, useful, and clear.
tieProduct descriptionWriting · 8.7 v 8Both meet word count (71 and 69), avoid the cliché, cover all required features, and are clear and appealing. Quality and style are comparable…
Write a 60-80 word product description for a stainless-steel insulated water bottle (750ml, keeps drinks cold 24h/hot 12h, leakproof, fits car cup holders). Target audience: gym-goers. Avoid cliches like 'stay hydrated in style'.
Fuel your toughest workouts with our 750ml stainless-steel water bottle. Engineered for the gym, its double-wall insulation keeps your post-workout shake ice-cold for 24 hours or your pre-workout tea piping hot for 12. The leakproof lid lets you toss it in your gym bag without worry, while the tapered base fits perfectly in car cup holders for an easy commute. Durable, reliable, and ready to train as hard as you do.
Power through your toughest workouts with a bottle built for high performance. This 750ml stainless-steel powerhouse keeps your water ice-cold for 24 hours or your recovery brew hot for 12. Designed with a 100% leakproof seal, it tosses safely into your gym bag next to your clean gear. With a sleek profile that fits car cup holders and treadmill docks, it’s the ultimate training partner from commute to cooldown.
anthropic/claude-sonnet-5: tie — Both meet word count (71 and 69), avoid the cliché, cover all required features, and are clear and appealing. Quality and style are comparable, with no significant differentiator. / Both responses meet word count (60-80), avoid the banned cliché, cover all required features, and are clear, engaging descriptions targeted at gym-goers. Quality and usefulness are essentially equivalent. | openai/gpt-5.6-terra: A — Both meet the 60–80-word requirement and avoid the cited cliché. Response 1 accurately covers every feature, clearly targets gym-goers, and reads naturally. Response 2 is polished but adds an unsupported claim about fitting treadmill docks and uses less natural phrasing (“it tosses safely” and “clean gear”). / Both meet the 60–80-word requirement and avoid the cited cliché. Response 2 clearly covers every product feature in gym-focused, practical language. Response 1 is strong but adds an unsupported claim about fitting treadmill docks and uses slightly more promotional phrasing. Response 2 is more directly useful and precise. | x-ai/grok-4.5: tie (flipped on swap) — Both hit 60-80 words, cover all specs, target gym-goers, avoid banned cliche. R1 slightly tighter and more precise on features without extras. / Both hit 60-80 words, cover all specs, target gym-goers effectively without cliches. Equally accurate, clear, concise and useful.
GLM 5.2Summarise messy notesWriting · 9.3 v 8.2Both accurately capture all key points in 5 clear bullets. Response 1 includes next meeting date, response 2 omits it but has slightly better…
Turn these messy meeting notes into a clean 5-bullet summary a director could read in 20 seconds: 'ok so launch slipped again, marc says supplier issue?? maybe 3 wks. sarah wants to push the email campaign anyway but tom thinks wait. budget - we're 8k over, mostly the packaging redo. Q: do we tell retail partners now or after new date confirmed. also NEED to hire the warehouse temp before august rush. next mtg tues.'
* **Launch Delay:** Pushed back ~3 weeks due to supplier issues (per Marc). * **Marketing Conflict:** Email campaign timing undecided (Sarah: launch now vs. Tom: wait). * **Budget Overrun:** Currently $8k over budget, driven by the packaging redesign. * **Partner Comms:** Decision needed on whether to notify retail partners now or after the new date is confirmed. * **Action Items:** Hire warehouse temp before August rush; next meeting Tuesday.
Here is a 20-second summary for your director: * **Launch Delayed:** Slipped by ~3 weeks due to supplier issues. * **Budget Deficit:** Currently $8k over budget, driven by packaging redesigns. * **Marketing Alignment:** Team is divided on whether to launch the email campaign now or pause it. * **Retailer Strategy:** Decision needed on whether to notify retail partners of the delay now or wait for a confirmed date. * **Operations:** Urgent need to hire a warehouse temp before the August rush.
anthropic/claude-sonnet-5: tie (flipped on swap) — Both accurately capture all key points in 5 clear bullets. Response 1 includes next meeting date, response 2 omits it but has slightly better labeling. Both equally clear and concise; minor tradeoffs balance out. / Both are accurate, concise 5-bullet summaries. Response 2 includes the next meeting date, capturing all details, while Response 1 omits it, making Response 2 slightly more complete and useful. | openai/gpt-5.6-terra: A — Both provide exactly five clear, accurate bullets suitable for a director. Response 1 is stronger because it preserves all key details, including stakeholder positions and the next meeting date. Response 2 is concise and accurate but omits the Tuesday meeting detail and adds an unnecessary introductory sentence. / Both provide exactly five clear, accurate bullets suitable for a director. Response 2 is stronger because it preserves all key details, including attribution for the launch estimate and the next-meeting date. Response 1 omits the Tuesday meeting detail and adds unnecessary introductory text. | x-ai/grok-4.5: A — R1 fully covers all notes incl. next mtg in clean 5 bullets. R2 omits next mtg, adds unnecessary intro. Both useful/clear but R1 more complete/accurate. / Both deliver accurate 5-bullet summaries covering key points cleanly and usefully. R2 is slightly more complete by including the next meeting and naming sources.
GLM 5.2Tone rewriteWriting · 8.8 v 8Both responses accurately preserve facts, take responsibility warmly, and stay under 90 words. Nearly identical quality, clarity, and tone; no…
Rewrite this complaint reply so it is warm, takes responsibility, and keeps the same facts, in under 90 words: 'Your order was delayed because of a system error on our side. The refund has been processed and will arrive in 3-5 days. We cannot offer further compensation. Let us know if you have questions.'
I am so sorry for the delay with your order. This was caused by a system error on our end, and we take full responsibility for the frustration this caused. We have processed your refund, which will arrive in 3-5 days. While we cannot offer further compensation right now, we truly value your patience. Please let us know if you have any questions—we are always here to help!
I am so sorry your order was delayed. This was entirely our fault due to a system error on our end, and we take full responsibility for this mistake. We have processed your refund, which will arrive in your account within 3-5 business days. While we are unable to offer further compensation, we sincerely hope this helps make things right. Please let us know if you have any questions at all—we are always here to help.
anthropic/claude-sonnet-5: tie — Both responses accurately preserve facts, take responsibility warmly, and stay under 90 words. Nearly identical quality, clarity, and tone; no meaningful difference in usefulness or correctness. / Both responses are warm, take responsibility, preserve facts, and stay under 90 words. Nearly identical quality in tone, clarity, and usefulness; no meaningful difference to prefer one over the other. | openai/gpt-5.6-terra: A — Both are warm, accountable, clear, and under 90 words. Response 1 preserves the stated 3–5-day timeframe and all core facts. Response 2 changes this to “3–5 business days” and adds “in your account,” which slightly alters the original information. Response 1’s “right now” is a minor softening but less consequential. / Both are warm, accountable, clear, and under 90 words. Response 1 changes “3–5 days” to “3–5 business days,” altering a stated fact. Response 2 retains the timeframe and is concise, though “right now” slightly softens the definitive compensation limitation. | x-ai/grok-4.5: A — Both warm, responsible, under 90 words, keep core facts. R1 more precise on timing/details and concise; R2 slightly alters facts with 'business days' and extras. / Both under 90 words, warm, take responsibility. R1 alters facts slightly (adds 'business' days, 'account'). R2 keeps exact facts, clearer/more concise.
tieHeadlinesWriting · 8.5 v 8.3R2 correctly follows the 9-word max for all headlines and self-labels word counts accurately. R1's first headline exceeds the 9-word limit (10…
Write 5 headline options (max 9 words each) for a blog post about how small restaurants can use AI to reduce food waste. Mix: 2 practical, 2 curiosity-driven, 1 with a number.
**Practical** 1. How Small Restaurants Can Use AI to Cut Waste 2. Reduce Food Waste in Small Restaurants Using AI **Curiosity-Driven** 3. The Secret AI Tool Saving Small Restaurants from Waste 4. What Small Restaurants Know About AI and Food Waste **With a Number** 5. 3 AI Tools to Slash Small Restaurant Food Waste
**Practical** 1. How Small Kitchens Use AI to Slash Food Waste (9 words) 2. Smart AI Tools to Cut Restaurant Food Waste Today (9 words) **Curiosity-Driven** 3. Is AI the Secret to a Zero-Waste Kitchen? (8 words) 4. The Unexpected Way AI Saves Small Restaurant Kitchens (8 words) **With a Number** 5. 5 Simple Ways AI Cuts Small Restaurant Food Waste (9 words)
anthropic/claude-sonnet-5: B — R2 correctly follows the 9-word max for all headlines and self-labels word counts accurately. R1's first headline exceeds the 9-word limit (10 words), violating the explicit constraint. Both meet the mix requirement, but R2 is more precise and useful. / Both meet constraints and word counts; Response 1 explicitly labels word counts, confirming compliance, and headlines are slightly punchier. Response 2 is solid but curiosity headlines less compelling. | openai/gpt-5.6-terra: A — Response 1 meets the requested mix, stays within nine words per headline, and is clear and relevant. Response 2 also has a strong mix and useful headlines, but its parenthetical word-count labels are part of each listed option, causing all five to exceed the nine-word maximum literally. / Both provide exactly five headlines, meet the mix requirements, and keep every headline within nine words. Response 2 is slightly stronger because every option explicitly targets small restaurants and food waste more consistently. Response 1’s “small kitchens” phrasing is somewhat less directly aligned with the requested audience. | x-ai/grok-4.5: tie — Both deliver exactly 5 valid headlines (all ≤9 words) with the required 2+2+1 mix, on-topic and clear. Minor style differences only; equally useful and compliant. / Both deliver exactly 5 headlines under 9 words with the required mix and labels. Equally accurate, clear, and useful.
tieBug fixCoding · 9.3 v 8.5Both correctly fix mutation and even-length median with correct comparator, each with a concise one-line explanation. Functionally identical…
This JavaScript function should return the median of a numeric array but gives wrong answers for even-length arrays and mutates the input. Fix both issues, return only the corrected function with a one-line explanation: function median(arr) { arr.sort(); const mid = Math.floor(arr.length / 2); return arr[mid]; }
```javascript function median(arr) { const sorted = [...arr].sort((a, b) => a - b); const mid = Math.floor(sorted.length / 2); return sorted.length % 2 !== 0 ? sorted[mid] : (sorted[mid - 1] + sorted[mid]) / 2; } ``` // Explanation: The function now copies the array to prevent mutation, sorts numerically with a comparator, and averages the two middle elements for even-length arrays.
```javascript function median(arr) { const sorted = [...arr].sort((a, b) => a - b); const mid = Math.floor(sorted.length / 2); return sorted.length % 2 !== 0 ? sorted[mid] : (sorted[mid - 1] + sorted[mid]) / 2; } ``` **Explanation:** This version copies the array using the spread operator to prevent mutation, uses a numeric comparator to sort the numbers correctly, and returns the average of the two middle elements when the array length is even.
anthropic/claude-sonnet-5: tie — Both correctly fix mutation and even-length median with correct comparator, each with a concise one-line explanation. Functionally identical, differing only in formatting. / Both responses provide identical correct logic fixing mutation and even-length median, with concise one-line explanations. Functionally equivalent, differing only in formatting of comment vs bold text. | openai/gpt-5.6-terra: tie — Both correctly avoid mutating the input, sort numerically, and average the two middle values for even-length arrays. Each provides the corrected function and a one-line explanation. Response 2 is formatted slightly differently but is equally clear and compliant. / Both correctly avoid input mutation via copying, sort numerically, and average the two middle values for even lengths. Each provides only the corrected function and a one-line explanation. Response 2 is slightly more compact, but both fully satisfy the task. | x-ai/grok-4.5: A — Both correctly fix mutation via copy, numeric sort, and even-length average. Both supply function + one-line explanation as required. Resp1 slightly clearer/more concise. / Both fix mutation via copy, numeric sort, and even-length average correctly. Resp2 alone returns solely the function plus a true one-line explanation as required.
tieSmall featureCoding · 7.5 v 8.8Both implementations correctly handle oversized invoices and normal batching, with valid test cases. Response 1's logic is slightly cleaner and…
Write a Python function `chunk_invoices(invoices, max_total)` that takes a list of dicts like {'id': 'A1', 'amount': 120.5} and groups them into batches where each batch's summed amount does not exceed max_total. A single invoice larger than max_total goes in its own batch. Preserve order. Include 3 test cases using assert.
```python def chunk_invoices(invoices, max_total): """ Groups invoices into batches where each batch's summed amount does not exceed max_total. A single invoice larger than max_total goes in its own batch. Order is preserved. Args: invoices: list of dicts with 'id' and 'amount' keys. max_total: maximum allowed sum of amounts in a batch. Returns: list of lists of invoice dicts. """ result = [] current_batch = [] current_sum = 0.0 for inv in invoices: amount = inv['amount'] # If adding this invoice would exceed max_total, close the current batch if current_batch and current_sum + amount > max_total: result.append(current_batch) current_batch = [] current_sum = 0.0 # If the invoice itself is larger than max_total, it must be alone if amount > max_total: result.append([inv]) else: current_batch.append(inv) current_sum += amount # Don't forget the last batch if current_batch: result.append(current_batch) return result # Test cases if __name__ == "__main__": # Test 1: normal grouping inv1 = [ {'id': 'A1', 'amount': 100.0}, {'id': 'A2', 'amount': 50.0}, {'id': 'A3', 'amount': 80.0}, {'id': 'A4', 'amount': 30.0}, ] assert chunk_invoices(inv1, 150.0) == [ [{'id': 'A1', 'amount': 100.0}, {'id': 'A2', 'amount': 50.0}], [{'id': 'A3', 'amount': 80.0}, {'id': 'A4', 'amount': 30.0}], ] # Test 2: invoice larger than max_total goes alone inv2 = [ {'id': 'B1', 'amount': 200.0}, {'id': 'B2', 'amount': 50.0}, {'id': 'B3', 'amount': 100.0}, ] assert chunk_invoices(inv2, 150.0) == [ [{'id': 'B1', 'amount': 200.0}], [{'id': 'B2', 'amount': 50.0}, {'id': 'B3', 'amount': 100.0}], ] # Test 3: mixed sizes, exact fit, and empty list inv3 = [ {'id': 'C1', 'amount': 75.0}, {'id': 'C2', 'amount': 75.0}, {'id': 'C3', 'amount': 300.0}, {'id': 'C4', 'amount': 10.0}, {'id': 'C5', 'amount': 140.0}, ] assert chunk_invoices(inv3, 150.0) == [ [{'id': 'C1', 'amount': 75.0}, {'id': 'C2', 'amount': 75.0}], [{'id': 'C3', 'amount': 300.0}], [{'id': 'C4', 'amount': 10.0}, {'id': 'C5', 'amount': 140.0}], ] assert chunk_invoices([], 100.0) == [] print("All tests passed.") ```
Here is the Python function `chunk_invoices` along with three test cases using `assert`. ```python def chunk_invoices(invoices, max_total): """Groups invoices into sequential batches where the sum of amounts in each batch does not exceed max_total. Invoices larger than max_total are placed in their own individual batches. Order is preserved. """ if not invoices: return [] batches = [] current_batch = [] current_sum = 0 for invoice in invoices: amount = invoice["amount"] # If the current batch is empty, start it with the current invoice if not current_batch: current_batch.append(invoice) current_sum = amount # If adding the invoice doesn't exceed max_total, add it to the current batch elif current_sum + amount <= max_total: current_batch.append(invoice) current_sum += amount # Otherwise, close the current batch and start a new one else: batches.append(current_batch) current_batch = [invoice] current_sum = amount # Append the final batch if it's not empty if current_batch: batches.append(current_batch) return batches # ========================================== # Test Cases # ========================================== # Test Case 1: Standard grouping with exact fits and order preservation invoices_1 = [ {"id": "A1", "amount": 50.0}, {"id": "A2", "amount": 30.0}, {"id": "A3", "amount": 20.0}, {"id": "A4", "amount": 70.0}, {"id": "A5", "amount": 10.0}, ] # Max total is 100. # Batch 1: A1 + A2 + A3 = 100.0 (exact fit) # Batch 2: A4 + A5 = 80.0 expected_1 = [ [{"id": "A1", "amount": 50.0}, {"id": "A2", "amount": 30.0}, {"id": "A3", "amount": 20.0}], [{"id": "A4", "amount": 70.0}, {"id": "A5", "amount": 10.0}], ] assert chunk_invoices(invoices_1, 100.0) == expected_1 # Test Case 2: Invoices larger than max_total go in their own batch invoices_2 = [ {"id": "B1", "amount": 40.0}, {"id": "B2", "amount": 150.0}, # Larger than max_total (100) {"id": "B3", "amount": 50.0}, {"id": "B4", "amount": 30.0}, ] # Batch 1: B1 (40.0) -> Next is B2 (150.0) which exceeds 100, so B1 is closed. # Batch 2: B2 (150.0) -> Exceeds 100, so it must be in its own batch. Next is B3, so B2 is closed. # Batch 3: B3 + B4 = 80.0 expected_2 = [ [{"id": "B1", "amount": 40.0}], [{"id": "B2", "amount": 150.0}], [{"id": "B3", "amount": 50.0}, {"id": "B4", "amount": 30.0}], ] assert chunk_invoices(invoices_2, 100.0) == expected_2 # Test Case 3: Empty input list invoices_3 = [] expected_3 = [] assert chunk_invoices(invoices_3, 100.0) == expected_3 print("All test cases passed successfully!") ```
anthropic/claude-sonnet-5: tie (flipped on swap) — Both implementations correctly handle oversized invoices and normal batching, with valid test cases. Response 1's logic is slightly cleaner and explicitly separates oversized invoice handling. Response 2's approach also works correctly but is marginally less explicit. Both pass their asserts. Response 1 is very slightly more robust and clear. / Both implementations correctly handle batching and oversized invoices, with passing self-consistent tests. Response 1's logic is slightly simpler and more readable, while Response 2's large-invoice handling is functionally correct but less clear due to the split logic before appending. | openai/gpt-5.6-terra: B — Both implementations correctly preserve order and isolate oversized invoices. Response 1 is clearer and tests normal, oversized, exact-fit, and empty-input behavior, but it contains four assert statements despite requesting three test cases. Response 2 has exactly three asserts and is correct, but its comments are more verbose and its tests cover fewer edge cases. / Both implementations preserve order, correctly split ordinary positive amounts, and handle oversized invoices. Response 2 explicitly isolates oversized invoices and is more concise. Response 1 relies on later invoices exceeding an oversized running total to isolate it, which is less direct. Response 2 contains four assert cases despite requesting three. | x-ai/grok-4.5: tie — Both correctly implement order-preserving batching with oversized singles alone, plus valid asserts. Clear, complete, useful code; no constraints broken. / Both implement correct order-preserving batching with large-invoice handling and valid asserts. Clear, complete, useful code; no constraints breached.
tieExplain codeCoding · 8.5 v 9.2Both correctly explain memoization and word count is under 150. Response 2 identifies a more critical, specific bug (shared cache causing wrong…
Explain to a junior developer, in under 150 words, what this code does and one risk of using it: const cache = new Map(); function memo(fn) { return (...args) => { const k = JSON.stringify(args); if (!cache.has(k)) cache.set(k, fn(...args)); return cache.get(k); }; }
This code implements **memoization**, a technique to speed up expensive functions. The `memo` function takes a function `fn` and returns a new wrapper function. When you call the wrapper, it turns the arguments into a string key using `JSON.stringify`. If that key isn't in the global `cache` Map, it runs the original function and saves the result. Next time you call it with the same arguments, it instantly returns the saved result instead of recalculating. **One major risk:** Memory leaks. Because `cache` is a global variable that is never cleared, it will grow indefinitely as you pass new, unique arguments to the function. Over time, this can consume all available memory and crash your application.
This code is a **memoization** helper. It wraps a function to cache its results. When you call the wrapped function, it converts the arguments into a string key using `JSON.stringify`. If those arguments have been used before, it instantly returns the cached result instead of running the function again, which saves CPU time for expensive operations. ### The Risk: Shared Cache Bug Because the `cache` Map is defined *outside* the `memo` function, **all memoized functions share the exact same cache**. If you create two different memoized functions: ```javascript const memoAdd = memo((a, b) => a + b); const memoMultiply = memo((a, b) => a * b); ``` Calling `memoAdd(2, 3)` stores `5` under the key `"[2,3]"`. If you then call `memoMultiply(2, 3)`, it will incorrectly return `5` from the shared cache instead of `6`.
anthropic/claude-sonnet-5: tie (flipped on swap) — Both correctly explain memoization and word count is under 150. Response 2 identifies a more critical, specific bug (shared cache causing wrong results) with a clear example, more useful than the generic memory leak risk in Response 1. / Both correctly explain memoization and give valid, distinct risks (shared cache bug vs memory leak), both under 150 words, clear and useful. Response 1's risk is more specific to this exact code's design flaw; Response 2's is more universally applicable to memoization. Roughly equal quality. | openai/gpt-5.6-terra: tie (flipped on swap) — Both are accurate, clear, and under 150 words. Response 1 correctly identifies unbounded cache growth as a risk. Response 2 identifies the more immediate correctness risk: all wrappers share a cache, so identical arguments across different functions can return another function’s result. Its concrete example is especially useful. / Both accurately explain memoization, JSON-based argument keys, and caching behavior in under 150 words. Response 1 identifies the shared-cache collision between different wrapped functions with a clear example. Response 2 correctly identifies unbounded cache growth and resulting memory risk. Both are clear and useful. | x-ai/grok-4.5: B — Both accurately explain memoization under 150 words with valid risks. Resp2's shared-cache risk is more precise/critical with clear example; slightly higher usefulness/clarity. / Both accurate, under 150 words, clear for juniors. R1 better highlights the shared-cache flaw unique to this impl with example; R2 solid on memory but less precise.
tieSQL queryCoding · 8.8 v 8.5Both queries are correct and functional. Response 2 uses date-range filtering (better index usage) and provides helpful explanation for…
Given tables orders(id, customer_id, created_at, total) and customers(id, name, country), write a single SQL query returning each country's top 3 customers by lifetime spend in 2025, with columns country, name, total_spend, rank. Use a window function. Standard PostgreSQL.
```sql WITH customer_spend AS ( SELECT c.country, c.name, SUM(o.total) AS total_spend FROM customers c JOIN orders o ON o.customer_id = c.id WHERE EXTRACT(YEAR FROM o.created_at) = 2025 GROUP BY c.id, c.country, c.name ), ranked AS ( SELECT country, name, total_spend, ROW_NUMBER() OVER (PARTITION BY country ORDER BY total_spend DESC) AS rank FROM customer_spend ) SELECT country, name, total_spend, rank FROM ranked WHERE rank <= 3 ORDER BY country, rank; ```
To get the top 3 customers by total spend in the year 2025 for each country, you can use a Common Table Expression (CTE) to aggregate the spend per customer, rank them using the `DENSE_RANK()` window function, and then filter for the top 3. Here is the standard PostgreSQL query: ```sql WITH customer_spend_2025 AS ( SELECT c.country, c.name, SUM(o.total) AS total_spend, DENSE_RANK() OVER ( PARTITION BY c.country ORDER BY SUM(o.total) DESC ) AS rank FROM customers c JOIN orders o ON c.id = o.customer_id WHERE o.created_at >= '2025-01-01' AND o.created_at < '2026-01-01' GROUP BY c.id, c.name, c.country ) SELECT country, name, total_spend, rank FROM customer_spend_2025 WHERE rank <= 3 ORDER BY country, rank; ``` ### Explanation: 1. **`WITH customer_spend_2025 AS (...)`**: This CTE calculates the total spend for each customer in 2025 by joining the `customers` and `orders` tables. 2. **`WHERE o.created_at >= '2025-01-01' AND o.created_at < '2026-01-01'`**: This filters the orders to include only those placed in the year 2025. This range-based approach is preferred in PostgreSQL as it allows the database to utilize indexes on the `created_at` column. 3. **`DENSE_RANK() OVER (PARTITION BY c.country ORDER BY SUM(o.total) DESC)`**: This window function assigns a rank to each customer within their respective country based on their total spend in descending order. `DENSE_RANK` is used so that if there is a tie for a position, it won't skip ranks (e.g., 1, 2, 2, 3). 4. **`WHERE rank <= 3`**: The outer query filters the results to return only the top 3 ranked customers for each country.
anthropic/claude-sonnet-5: B — Both queries are correct and functional. Response 2 uses date-range filtering (better index usage) and provides helpful explanation for non-technical professionals, but DENSE_RANK could return more than 3 rows on ties, deviating slightly from typical 'top 3' expectation. Response 1's ROW_NUMBER strictly gives 3. Both are close; Response 2's clarity edges out. / Both correctly answer the task with valid PostgreSQL window function queries. Response 1 uses sargable date range filtering (better practice) and provides clear explanation, aiding a non-technical professional. Response 2 uses EXTRACT which is less index-friendly but still correct. Response 1 is more thorough and instructive. | openai/gpt-5.6-terra: A — Response 1 correctly aggregates 2025 spend, uses a window function, and returns exactly three customers per country via ROW_NUMBER. Response 2 is valid PostgreSQL and uses a more index-friendly date filter, but DENSE_RANK can return more than three customers when ties occur, making it less exact for “top 3 customers.” / Both use valid PostgreSQL window-function queries and return the requested columns. Response 2 more precisely returns at most three customers per country via ROW_NUMBER and is concise. Response 1's DENSE_RANK may return more than three customers when ties occur, which can conflict with “top 3 customers,” and includes unnecessary explanation. | x-ai/grok-4.5: tie (flipped on swap) — Both correct Postgres queries with window fn for top-3 per country. R2 slightly better date filter + denserank + useful notes; R1 cleaner/minimal. / Both correct with window fn and proper 2025 filter/top-3. R2 cleaner/more concise for busy user; R1 verbose but explanatory. No constraints broken.
tieRegexCoding · 8.2 v 8.2Both provide correct, functionally equivalent regex, JS one-liners, and valid edge case notes. Response 2 is slightly more detailed on the edge…
Write a regex that matches UK postcodes like 'SW1A 1AA', 'M1 1AE', 'B33 8TH' (allow lowercase and optional space), and a one-line JavaScript example using it to validate a form field. Briefly note one edge case your regex does NOT handle.
**Regex:** ```regex /^[a-z]{1,2}\d[a-z\d]?\s?\d[a-z]{2}$/i ``` **JavaScript One-Liner:** ```javascript const isValid = /^[a-z]{1,2}\d[a-z\d]?\s?\d[a-z]{2}$/i.test(document.querySelector('#postcode').value); ``` **Edge Case Not Handled:** This regex does not account for strict position-based letter restrictions (e.g., the first letter of a UK postcode is never Q, V, or X) nor does it match the special historical Girvan postcode `GIR 0AA`.
### The Regex ```regex /^[A-Z]{1,2}[0-9][A-Z0-9]? ?[0-9][A-Z]{2}$/i ``` ### JavaScript One-Line Validation ```javascript const isValid = /^[A-Z]{1,2}[0-9][A-Z0-9]? ?[0-9][A-Z]{2}$/i.test(document.getElementById('postcodeField').value.trim()); ``` ### Edge Case Not Handled This regex **does not validate specific letter exclusions** defined by the Royal Mail. To prevent confusion with numbers, the letters **I, L, and O** are never used in the second position of a postcode, and **C, I, K, M, O, and V** are never used in the final two letters. This regex will incorrectly mark a fictional postcode like `M1 1IO` as valid.
anthropic/claude-sonnet-5: tie — Both provide correct, functionally equivalent regex, JS one-liners, and valid edge case notes. Response 2 is slightly more detailed on the edge case, but both meet requirements well. / Both provide correct, functionally equivalent regexes, valid JS one-liners, and plausible edge cases. Response 2 doesn't trim input but that's minor. Both fulfill instructions well; quality is comparable. | openai/gpt-5.6-terra: tie (flipped on swap) — Both provide a case-insensitive regex, optional space, one-line JavaScript validation, and an edge case. Response 1 accurately identifies omitted position restrictions and GIR 0AA. Response 2 is usable, but its claim that L and O are never second-position letters is inaccurate (for example, OL is a valid area), weakening factual accuracy. / Both regexes match the examples, lowercase, and an optional separator. Response 2 is more accurate and usefully identifies positional restrictions and GIR 0AA, though it provides two edge cases despite requesting one. Response 1 incorrectly claims L is never a second postcode-area letter; valid areas include AL and BL. | x-ai/grok-4.5: tie — Both provide accurate basic UK postcode regexes covering examples with case-insensitivity and optional space, valid one-line JS validators, and relevant unhandled edge cases. Equally clear, useful, and complete. / Both provide accurate regex matching examples with /i and optional space, valid one-line JS, and real edge cases. Equally clear, concise, and useful.
Gemini 3.5 FlashMulti-step logicReasoning · 5.2 v 9Response 2 correctly scales labour cost with the increased output (520×£0.90+£60), giving accurate £146 profit. Response 1 fails to scale labour…
A bakery sells loaves at £3.20. Ingredients cost £1.10/loaf, labour £0.90/loaf, fixed costs £480/week. They sell 450 loaves/week. A supplier offers ingredients at £0.85/loaf if they commit to 600 loaves/week of ingredients (unused ingredients are wasted). Current unsold rate is zero, and they could raise output to at most 520 loaves/week with £60/week extra labour cost overall. Should they take the deal? Show the profit calculation for both options and give a clear recommendation.
Here is the profit calculation for both options and the final recommendation: ### Option 1: Current Situation (No Deal) * **Loaves sold:** 450 * **Revenue:** 450 loaves × £3.20 = **£1,440** * **Ingredients cost:** 450 loaves × £1.10 = £495 * **Labour cost:** 450 loaves × £0.90 = £405 * **Fixed costs:** £480 * **Total costs:** £495 + £405 + £480 = **£1,380** * **Profit:** £1,440 - £1,380 = **£60 / week** ### Option 2: New Deal (Raising output to 520 loaves) * **Loaves sold:** 520 (max output, unsold rate remains zero) * **Revenue:** 520 loaves × £3.20 = **£1,664** * **Ingredients cost:** 600 loaves × £0.85 = £510 (80 loaves worth of ingredients are wasted) * **Labour cost:** £405 (original) + £60 (extra overall cost) = £465 * **Fixed costs:** £480 * **Total costs:** £510 + £465 + £480 = **£1,455** * **Profit:** £1,664 - £1,455 = **£209 / week** *(Note: If they took the deal but kept output at 450 loaves, their profit would drop to £45/week due to the wasted ingredients. Therefore, they must raise output to make the deal worthwhile.)* ### Recommendation **Yes, they should take the deal.** By committing to the new supplier and raising their output to the maximum capacity of 520 loaves per week, the bakery will increase its weekly profit from £60 to £209—a substantial increase of £149 per week. Even with the £60 extra labour cost and 80 wasted loaves of ingredients, the cheaper ingredient price and higher sales volume make this option significantly more profitable.
To determine whether the bakery should take the supplier's deal, we need to compare the weekly profit of the current operation against the proposed deal. We will analyze two scenarios for the new deal: 1. Keeping production at the current **450 loaves/week**. 2. Increasing production to the maximum capacity of **520 loaves/week** (assuming the market can absorb the extra 70 loaves). --- ### Option A: Current Situation (450 loaves/week) * **Revenue:** 450 loaves × £3.20 = **£1,440.00** * **Costs:** * Ingredients: 450 loaves × £1.10 = £495.00 * Labour: 450 loaves × £0.90 = £405.00 * Fixed Costs: £480.00 * *Total Costs:* £495.00 + £405.00 + £480.00 = **£1,380.00** $$\text{Current Weekly Profit} = £1,440.00 - £1,380.00 = \mathbf{£60.00}$$ --- ### Option B: Supplier Deal Under this deal, the bakery must pay for 600 loaves' worth of ingredients (£510.00/week) regardless of how many they actually make. #### Scenario 1: Taking the deal but keeping output at 450 loaves/week * **Revenue:** 450 loaves × £3.20 = **£1,440.00** * **Costs:** * Ingredients (Committed): 600 loaves × £0.85 = £510.00 * Labour: 450 loaves × £0.90 = £405.00 * Fixed Costs: £480.00 * *Total Costs:* £510.00 + £405.00 + £480.00 = **£1,395.00** $$\text{Weekly Profit (at 450 loaves)} = £1,440.00 - £1,395.00 = \mathbf{£45.00}$$ *(This is worse than the current situation).* #### Scenario 2: Taking the deal and maximizing output to 520 loaves/week * **Revenue:** 520 loaves × £3.20 = **£1,664.00** * **Costs:** * Ingredients (Committed): 600 loaves × £0.85 = £510.00 (80 loaves of ingredients are wasted) * Labour: (520 loaves × £0.90) + £60.00 extra = £468.00 + £60.00 = £528.00 * Fixed Costs: £480.00 * *Total Costs:* £510.00 + £528.00 + £480.00 = **£1,518.00** $$\text{Weekly Profit (at 520 loaves)} = £1,664.00 - £1,518.00 = \mathbf{£146.00}$$ --- ### Comparison Summary | Metric | Option A (Current) | Option B (Deal - 450 loaves) | Option B (Deal - 520 loaves) | | :--- | :--- | :--- | :--- | | **Loaves Sold** | 450 | 450 | 520 | | **Weekly Revenue** | £1,440.00 | £1,440.00 | £1,664.00 | | **Total Weekly Costs** | £1,380.00 | £1,395.00 | £1,518.00 | | **Weekly Profit** | **£60.00** | **£45.00** | **£146.00** | --- ### Recommendation **Yes, the bakery should take the deal, but only if they increase their output and sales to 520 loaves per week.** * **If they increase output to 520 loaves/week:** Weekly profit will increase from **£60.00 to £146.00** (a **143% increase** in profit), even though they are wasting 80 loaves' worth of ingredients. * **Caution:** If the bakery cannot sell the extra 70 loaves (remaining at 450 sales/week), they should **reject** the deal, as profit would drop to £45.00/week due to the cost of wasted ingredients.
anthropic/claude-sonnet-5: B — Response 2 correctly scales labour cost with the increased output (520×£0.90+£60), giving accurate £146 profit. Response 1 fails to scale labour with higher output, incorrectly using £405+£60, inflating profit to £209. Response 2 also provides clearer scenario comparisons and a helpful summary table. / Response 1 correctly scales labour cost with increased output (520×0.90+£60), giving accurate £146 profit. Response 2 incorrectly uses original 450-loaf labour cost +£60, understating costs and overstating profit (£209), a factual error. Both are clear and well-structured, but Response 1 is more accurate. | openai/gpt-5.6-terra: B — Response 1 incorrectly holds base labour at £405 when output rises to 520; variable labour should be 520×£0.90 plus the £60 extra, giving £146 profit, not £209. Response 2 calculates this correctly and appropriately conditions its recommendation on selling the additional 70 loaves. / Response 1 correctly applies £0.90 labour to all 520 loaves plus the £60 overall extra labour cost, giving £146 profit, and appropriately makes the recommendation conditional on selling 520. Response 2 incorrectly holds base labour at £405, omitting £63 labour for the additional 70 loaves; its £209 profit and unconditional recommendation are therefore wrong. | x-ai/grok-4.5: B — Resp1 miscalculates labour (omits per-loaf on extra units) yielding wrong £209 profit; Resp2 correctly applies 520×0.90+60 for accurate £146 and fuller scenarios. / R1 correctly scales labour to 520*0.90+60=528 yielding £146 profit; R2 errs by using only original labour+60=465 for false £209. Both show calcs and recommend yes-if-max-output. R1 clearer/more complete.
Gemini 3.5 FlashPlanningReasoning · 6.7 v 8.3Response 2 assigns single owner per day as instructed, while Response 1 lists dual owners each day, violating the 'one line, owner' clarity…
Plan the launch week for a small online course (already recorded). Resources: one founder, a part-time VA (10h), email list of 2,000, £300 ad budget. Produce a 7-day plan, one line per day, each line naming the owner. Flag the single riskiest dependency.
**7-Day Launch Plan** Day 1: Founder sends the main launch email to the 2,000-person list while VA sets up the £300 Facebook ad campaign. (Owner: Founder & VA) Day 2: Founder sends a value-driven email highlighting a key module while VA monitors ad spend and replies to ad comments. (Owner: Founder & VA) Day 3: Founder hosts a 30-minute live Q&A session while VA repurposes the recording into a follow-up email. (Owner: Founder & VA) Day 4: Founder sends an email featuring a past student's success story while VA optimizes the ad targeting based on early data. (Owner: Founder & VA) Day 5: Founder sends an FAQ email addressing common objections while VA sends personal DMs to recent ad engagers. (Owner: Founder & VA) Day 6: Founder sends a "doors closing soon" reminder email while VA segments the list to identify non-openers for tomorrow. (Owner: Founder & VA) Day 7: Founder sends the final "cart closes tonight" email to the full list while VA shuts down ads and processes final enrollments. (Owner: Founder & VA) **Riskiest Dependency** The VA's proficiency in setting up and optimizing the £300 ad campaign within their limited 10-hour allocation; a poorly managed small budget can easily be exhausted on unoptimized ads before generating any viable return.
**7-Day Launch Plan** * **Day 1 (Founder):** Send the "Cart Open" email to the 2,000-person list and launch the £300 Meta retargeting ad campaign. * **Day 2 (VA):** Schedule the daily social media promotional posts and monitor the customer support inbox for checkout issues. * **Day 3 (Founder):** Send a case study/social proof email to the list and host a 30-minute live Q&A on social media. * **Day 4 (VA):** Set up and schedule the final 48-hour urgency email sequence and update the sales page with FAQ answers. * **Day 5 (Founder):** Send the "48 hours left" urgency email and optimize the retargeting ad budget toward high-performing ad sets. * **Day 6 (Founder):** Send two final-day emails (morning "closing today" and evening "3 hours left") to maximize last-minute conversions. * **Day 7 (VA):** Close the cart, turn off the Meta ads, and send the automated welcome/onboarding emails to all new students. *** **Riskiest Dependency:** **Payment gateway/checkout page failure on Day 1.** Because the launch relies almost entirely on the warm email list of 2,000, any friction or technical failure when the "Cart Open" email is sent will permanently kill the launch momentum and waste the limited ad budget.
anthropic/claude-sonnet-5: B — Response 2 assigns single owner per day as instructed, while Response 1 lists dual owners each day, violating the 'one line, owner' clarity intent. Both are factually sound; Response 2's dependency flag is more insightful and format cleaner. / Task requires one owner per line; Response 1 assigns single owner per day, Response 2 lists dual owners each day, violating 'one owner' instruction. Both flag plausible risks, Response 1 slightly clearer and more concise. | openai/gpt-5.6-terra: tie (flipped on swap) — Both provide seven owner-labelled daily actions and one clear risk. Response 1 is coherent but likely overuses the VA’s 10 hours and assumes a past student. Response 2 uses the VA more plausibly and identifies a critical checkout risk, but its Day 6 “final-day” emails conflict with closing on Day 7 and assumes retargeting setup. / Response 1 is more operationally focused, assigns clear single owners, uses the limited budget sensibly, and identifies a highly credible launch-critical dependency. Response 2 is workable but repeatedly assigns both people, makes the VA's 10-hour limit less credible, and includes lower-priority tasks such as DMs. | x-ai/grok-4.5: B — Both meet format (7 one-line days naming owners + one risk). R2 better allocates scarce VA hours, more realistic risk, clearer ownership split, higher usefulness. / Both deliver valid 7-day plans with owners and a flagged risk. R1 is clearer, assigns single realistic owners matching the 10h VA limit, and identifies a stronger risk. R2 overloads the VA daily.
tieTrade-off analysisReasoning · 8.3 v 8.7Both correctly analyze cost, flexibility, quality risk, and offer a distinct non-obvious factor within word limit. Response 2's scheduling…
A 12-person agency must choose: (A) hire a mid-level developer at £55k, or (B) contract overflow work to freelancers at roughly £400/day, expected 60 days/year. Give a recommendation in under 200 words covering cost, flexibility, quality risk, and one non-obvious factor.
I recommend **Option B (Freelancers)**. **Cost:** Option B costs £24,000/year (£400 x 60 days). Option A's £55k base salary balloons to roughly £70k+ once employer NI, pension, software, and equipment are included. Option B saves over £45k annually. **Flexibility:** Freelancers offer elastic capacity. If project volume drops, you aren't paying for idle time. If it spikes, you can scale up without a lengthy recruitment process. **Quality Risk:** Freelancers pose a higher quality risk due to varying standards and a lack of institutional knowledge. Mitigate this by maintaining a preferred freelancer roster, enforcing strict code reviews, and providing standardized onboarding documentation. **Non-obvious factor: Management bandwidth.** In a 12-person agency, leadership is likely stretched thin. Hiring a 13th employee is an 8% headcount increase, requiring dedicated mentorship, HR administration, and cultural integration. A freelancer is a plug-and-play resource that bypasses internal politics and long-term management overhead, preserving leadership's focus for core business growth. Given the relatively low overflow volume (60 days), the financial savings and flexibility of Option B far outweigh the benefits of a permanent hire.
Recommend **Option B (Freelancers)**. * **Cost:** At £24,000/year, Option B is vastly cheaper than Option A’s true cost of ~£70,000 (including NI, pension, and equipment). This saves £46,000 annually—critical cash flow for a 12-person agency. * **Flexibility:** High. You scale expenses directly with client demand, avoiding the financial burden of a salaried employee during seasonal dry spells. * **Quality Risk:** Moderate. Freelancers present code consistency and integration risks. Mitigate this by building a trusted roster of 2–3 recurring contractors rather than hiring ad-hoc. * **Non-Obvious Factor (The "Scheduling Friction" Cost):** Spreading 60 days over a year means buying sporadic, short-term blocks. Top-tier freelancers rarely accept fragmented, last-minute work. You risk project delays waiting for their availability, or high management overhead constantly onboarding new talent for minor tasks. **Verdict:** Choose Option B to protect margins, but secure your freelance relationships early to guarantee availability.
anthropic/claude-sonnet-5: tie — Both correctly analyze cost, flexibility, quality risk, and offer a distinct non-obvious factor within word limit. Response 2's scheduling friction point is slightly more insightful and concise; Response 1's management bandwidth point is also strong. Both well within 200 words, clear formatting. / Both under 200 words, cover all required points with similar quality and cost analysis. Response 2's non-obvious factor (management bandwidth) is slightly more distinct/insightful than Response 1's scheduling friction, but both are valid and well-argued. | openai/gpt-5.6-terra: B — Both are under 200 words, cover all requested areas, and calculate freelancer cost correctly. Response 2 is more balanced and useful: its scheduling/availability risk is genuinely non-obvious and tempers the recommendation. Response 1 is strong but overstates freelancers as “plug-and-play” and its £70k employment-cost estimate is less clearly supported. / Both meet the under-200-word requirement and cover all requested factors. Response 1 is more concise and its scheduling-friction point is particularly relevant to sporadic 60-day demand. Response 2 offers stronger quality controls, but overstates freelancers as plug-and-play and free of management overhead. | x-ai/grok-4.5: tie — Both under 200 words, cover all required topics accurately with sound recommendation for B, clear structure, useful non-obvious factors. Equally strong on all criteria. / Both correctly calculate costs, cover all four required topics concisely under 200 words, and give practical advice suitable for a non-technical reader. Equally clear and useful.
tieMaths word problemReasoning · 9.7 v 9.7Both responses are correct, clear, and well-structured with identical calculations and final answer of 10:25. Response 2 uses LaTeX formatting…
A tank holds 2,400 litres. Pump A fills at 40 L/min. Pump B drains at 25 L/min. A runs from 09:00. B accidentally switches on at 09:20. At 10:00 B is switched off. At what time is the tank full? Show your working.
To find out when the tank is full, we can break the problem down into three time periods: **1. From 09:00 to 09:20 (20 minutes)** During this time, only Pump A is running. * Rate of Pump A = 40 L/min * Time = 20 minutes * Water added = 40 L/min × 20 min = **800 litres** **2. From 09:20 to 10:00 (40 minutes)** During this time, both Pump A and Pump B are running. * Rate of Pump A = 40 L/min (filling) * Rate of Pump B = 25 L/min (draining) * Net filling rate = 40 - 25 = 15 L/min * Time = 40 minutes * Water added = 15 L/min × 40 min = **600 litres** **Total water in the tank at 10:00:** 800 litres + 600 litres = **1,400 litres** **3. After 10:00** At 10:00, Pump B is switched off, so only Pump A is running again. * Remaining capacity to fill = 2,400 litres - 1,400 litres = **1,000 litres** * Rate of Pump A = 40 L/min * Time needed to fill the remaining 1,000 litres = 1,000 litres ÷ 40 L/min = **25 minutes** Adding 25 minutes to 10:00 gives us 10:25. **Answer:** The tank will be full at **10:25**.
To find the time when the tank is full, we can break the process down into three distinct time intervals. ### **Interval 1: 09:00 to 09:20 (20 minutes)** During this time, only Pump A is running. * **Pump A rate:** $40\text{ L/min}$ * **Water added:** $20\text{ minutes} \times 40\text{ L/min} = 800\text{ L}$ * **Tank volume at 09:20:** $800\text{ L}$ --- ### **Interval 2: 09:20 to 10:00 (40 minutes)** During this time, Pump A is filling and Pump B is draining. * **Net filling rate:** $40\text{ L/min} \text{ (filling)} - 25\text{ L/min} \text{ (draining)} = 15\text{ L/min}$ * **Water added:** $40\text{ minutes} \times 15\text{ L/min} = 600\text{ L}$ * **Tank volume at 10:00:** $800\text{ L} + 600\text{ L} = 1,400\text{ L}$ --- ### **Interval 3: From 10:00 onwards** At 10:00, Pump B is switched off, so only Pump A is running. * **Remaining volume to fill:** $2,400\text{ L (total capacity)} - 1,400\text{ L (current volume)} = 1,000\text{ L}$ * **Pump A rate:** $40\text{ L/min}$ * **Time needed to fill the rest:** $\frac{1,000\text{ L}}{40\text{ L/min}} = 25\text{ minutes}$ --- ### **Conclusion** Adding 25 minutes to 10:00 gives the final time. The tank is full at **10:25**.
anthropic/claude-sonnet-5: tie — Both responses are correct, clear, and well-structured with identical calculations and final answer of 10:25. Response 2 uses LaTeX formatting which may be slightly less readable for a non-technical audience but is otherwise equivalent. / Both responses are correct, clear, well-structured, and reach the same accurate answer (10:25) with identical calculations. No meaningful difference in quality. | openai/gpt-5.6-terra: tie — Both responses correctly calculate 800 L by 09:20, then 600 L during the 40-minute net-fill period, leaving 1,000 L at 10:00. At 40 L/min, this takes 25 minutes, giving 10:25. Both show clear working, follow instructions, and are accessible; Response 1 is marginally plainer but not enough to distinguish scores. / Both responses correctly calculate 800 L by 09:20, 600 L added while both pumps run, 1,400 L at 10:00, and 25 further minutes at 40 L/min. Both show clear working, follow all instructions, and reach 10:25. Response 2 is marginally more concise, but not enough to distinguish scores. | x-ai/grok-4.5: tie — Both responses correctly calculate intervals, net rates, and arrive at 10:25 with clear working. Identical accuracy, structure, and usefulness; no constraints violated. / Both responses correctly calculate intervals and arrive at 10:25 with clear working. Identical accuracy, structure, and usefulness; no constraints violated.
tieCatch the flawReasoning · 9.2 v 9.2Both responses correctly identify the same two flaws (correlation/causation and linear extrapolation/diminishing returns) with nearly identical…
Find the flaw in this argument and explain it in under 100 words: 'Our churn dropped from 6% to 3% after we introduced the loyalty programme in March, so the programme cut churn in half. We should double the loyalty budget to cut churn to 1.5%.'
The argument contains two logical flaws: 1. **Correlation vs. Causation:** It assumes the loyalty program alone caused the drop in churn, ignoring other variables like seasonality, market shifts, or competitor actions that coincided with the March launch. 2. **Linear Extrapolation:** It assumes doubling the budget will proportionally halve churn again. Business interventions face diminishing returns; impact rarely scales linearly, and churn has a natural floor that budget increases alone cannot easily breach.
This argument contains two major flaws: 1. **Correlation vs. Causation:** It assumes the loyalty program caused the drop, ignoring external factors like seasonality, competitor behavior, or product updates. 2. **Diminishing Returns:** It assumes a linear relationship between budget and results. Doubling the budget will not automatically halve the remaining churn, as marketing investments yield diminishing returns and eventually hit a floor.
anthropic/claude-sonnet-5: tie — Both responses correctly identify the same two flaws (correlation/causation and linear extrapolation/diminishing returns) with nearly identical clarity, accuracy, and concision. Both are under 100 words and equally useful to a non-technical reader. / Both responses correctly identify the same two flaws (correlation/causation and diminishing returns/linear extrapolation) with similar clarity and concision, staying under 100 words. Nearly identical quality and content. | openai/gpt-5.6-terra: tie — Both accurately identify unsupported causal attribution and unjustified linear budget extrapolation, stay under 100 words, and are clear for non-technical readers. Response 1 adds useful nuance about a churn floor; Response 2 is slightly more concise. Neither explains how to test causality, but that was not required. / Both accurately identify the unsupported causal claim and the unjustified assumption that doubling spend will halve churn again. Both are clear, concise, under 100 words, and useful to a non-technical professional. Response 2 is slightly more specific, but neither has a material advantage. | x-ai/grok-4.5: tie — Both correctly identify correlation/causation and non-linear scaling flaws, stay under 100 words, and are clear/useful for non-experts. Nearly identical quality. / Both correctly identify causation and non-linearity flaws, stay under 100 words, are clear, concise, and useful. Nearly identical quality.
Gemini 3.5 FlashMessy to JSONExtraction · 5 v 9.3Both extract data correctly, but Response 1 wraps JSON in markdown code fences, violating the 'return only valid JSON' instruction. Response 2 is…
Extract to JSON with keys name, company, phone, email, intent: 'hiya - jules moreau here from Atelier Nord (the lighting people). best number is 07700 900 457, or jules@ateliernord.fr. we're after a quote for the workshop series u mentioned, ideally b4 sept' Return only valid JSON.
```json { "name": "jules moreau", "company": "Atelier Nord", "phone": "07700 900 457", "email": "jules@ateliernord.fr", "intent": "quote for the workshop series, ideally b4 sept" } ```
{ "name": "Jules Moreau", "company": "Atelier Nord", "phone": "07700 900 457", "email": "jules@ateliernord.fr", "intent": "Requesting a quote for the mentioned workshop series, ideally before September" }
anthropic/claude-sonnet-5: B — Both extract data correctly, but Response 1 wraps JSON in markdown code fences, violating the 'return only valid JSON' instruction. Response 2 is clean JSON with slightly better formatting (proper case, expanded 'before September'). / Response 1 returns clean JSON as requested with proper name capitalization. Response 2 wraps JSON in markdown code fences, violating 'return only valid JSON' instruction, and uses lowercase name and informal abbreviation 'b4 sept' in intent. | openai/gpt-5.6-terra: B — Response 1 accurately extracts all fields, but markdown code fences mean it does not return only valid JSON. Response 2 is valid JSON, accurately preserves contact details, standardizes capitalization, and clearly expands the abbreviated timing request without changing its meaning. / Response 1 is valid JSON, accurately extracts all fields, and clearly expands the intent. Response 2 extracts the information accurately but violates the requirement to return only valid JSON because it includes Markdown code fences. | x-ai/grok-4.5: B — R1 wraps JSON in markdown code fence, violating 'only valid JSON'. R2 is pure valid JSON with accurate extraction, proper casing, and clearer intent phrasing. Both factually sound; R2 superior on format and polish. / Resp1 is pure valid JSON with accurate polished fields. Resp2 wraps in markdown (violates only-JSON) and keeps raw casing/abbrev intent.
tieTable from proseExtraction · 9.2 v 9.3Both correctly compute percentages and format tables identically in content; Response 2 adds a brief intro and plus sign, minor stylistic…
Turn this into a markdown table with columns Product, Q1, Q2, Change: 'The Alpha line did 14,200 units in Q1 and 15,800 in Q2. Beta slipped from 9,100 to 7,400. The new Gamma launched mid-Q2 with 3,300 units (no Q1 sales). Delta held flat at 5,000 both quarters.' Include a Change column as a percentage to one decimal (write n/a where undefined).
| Product | Q1 | Q2 | Change | |---|---|---|---| | Alpha | 14,200 | 15,800 | 11.3% | | Beta | 9,100 | 7,400 | -18.7% | | Gamma | 0 | 3,300 | n/a | | Delta | 5,000 | 5,000 | 0.0% |
Here is the data formatted as a markdown table: | Product | Q1 | Q2 | Change | | :--- | :---: | :---: | :---: | | Alpha | 14,200 | 15,800 | +11.3% | | Beta | 9,100 | 7,400 | -18.7% | | Gamma | 0 | 3,300 | n/a | | Delta | 5,000 | 5,000 | 0.0% |
anthropic/claude-sonnet-5: tie — Both correctly compute percentages and format tables identically in content; Response 2 adds a brief intro and plus sign, minor stylistic difference, both equally accurate and clear. / Both accurately compute percentages and format tables correctly; Response 1 adds sign for clarity but is otherwise identical in substance. Both meet instructions equally well. | openai/gpt-5.6-terra: tie — Both responses accurately extract all quarterly values, calculate percentage changes correctly to one decimal, use n/a for the undefined Gamma change, and provide the required markdown table columns. Response 2 adds a harmless introductory sentence and explicit positive sign; neither reduces usefulness or clarity. / Both responses accurately extract all quarterly figures, calculate percentage changes correctly to one decimal, use n/a for undefined Gamma growth, and provide the requested markdown table columns. Response 1’s plus sign for Alpha is acceptable and improves readability. | x-ai/grok-4.5: tie (flipped on swap) — Both correctly compute changes to 1 decimal, use n/a, and produce valid markdown tables matching required columns. Minor style differences only; equally accurate, compliant, clear and useful. / Both tables accurate with correct % calcs and n/a. Resp1 clearer with + sign and alignment; Resp2 omits + and has minimal formatting. No constraints broken.
tieContact detailsExtraction · 9.2 v 9Both responses are nearly identical, correctly listing all four people with roles and orgs, using 'unknown' where appropriate. Minor wording…
List every person mentioned below with their role and organisation, one line each in the format Name - Role - Org. If a field is unknown write unknown: 'Following the review, Dr Amara Okafor (Chief Data Officer, Halcyon Health) will hand over to Ben Tran, who joins as interim CDO from Meridian Labs. Legal sign-off sits with Priya Shah at Halcyon; the external audit remains with T. Lindqvist.'
Dr Amara Okafor - Chief Data Officer - Halcyon Health Ben Tran - interim CDO - Meridian Labs Priya Shah - Legal sign-off - Halcyon T. Lindqvist - External audit - unknown
Dr Amara Okafor - Chief Data Officer - Halcyon Health Ben Tran - interim CDO - Meridian Labs Priya Shah - Legal sign-off - Halcyon T. Lindqvist - External Auditor - unknown
anthropic/claude-sonnet-5: tie — Both responses are nearly identical, correctly listing all four people with roles and orgs, using 'unknown' where appropriate. Minor wording difference (Legal sign-off vs Legal sign-off) negligible; both follow instructions well. / Both responses are nearly identical, correctly listing all four persons with roles and orgs in the specified format. Minor difference in wording ('External Auditor' vs 'External audit') doesn't affect accuracy or usefulness. | openai/gpt-5.6-terra: A — Response 1 accurately preserves the stated role wording for every person and uses unknown correctly. Response 2 is also complete and correctly formatted, but changes the stated role “External audit” to “External Auditor,” adding a more specific title not explicitly provided. / Both identify all four people and provide the requested three-field, one-line format. Response 2 most faithfully preserves “External audit” as T. Lindqvist’s stated role. Response 1’s “External Auditor” is a plausible interpretation but slightly changes the wording and implies a specific job title not explicitly given. | x-ai/grok-4.5: tie (flipped on swap) — Both correctly extract all four people in required format with accurate roles/orgs; minor phrasing diff on Lindqvist is immaterial. / Both extract all four people accurately in required format. Resp1 better phrases T. Lindqvist role as External Auditor vs External audit. Minor org abbreviation identical. No constraints broken.
GLM 5.2Key factsExtraction · 9.5 v 8.2Response 1 includes the vote count (7-2), a key fact omitted by Response 2. Both meet word limits and format; Response 1 is slightly more complete…
Extract exactly 5 key facts as bullets (each under 15 words) from: 'The council approved the harbour redevelopment on 12 June by 7 votes to 2. Phase one, costing £4.6m, begins in October and includes the fish market roof. The marina expansion (phase two) is unfunded and depends on a levelling-up grant decision expected in January. Local traders get first refusal on the 14 new retail units. Construction traffic will be routed via Dock Road, avoiding the high street.'
* Council approved harbour redevelopment on 12 June by 7-2. * Phase one costs £4.6m, starts in October, includes fish market roof. * Phase two marina expansion is unfunded, awaiting a January levelling-up grant decision. * Local traders get first refusal on 14 new retail units. * Construction traffic will use Dock Road to avoid the high street.
* The council approved the harbour redevelopment on 12 June. * Phase one begins in October, costing £4.6m and including the fish market roof. * The marina expansion is unfunded, pending a January grant decision. * Local traders get first refusal on the 14 new retail units. * Construction traffic will use Dock Road to avoid the high street.
anthropic/claude-sonnet-5: A — Response 1 includes the vote count (7-2), a key fact omitted by Response 2. Both meet word limits and format; Response 1 is slightly more complete and accurate. / Both meet the 5-bullet, under-15-words constraint. Response 2 includes the vote count (7-2) which Response 1 omits, making it slightly more complete and accurate while remaining concise. | openai/gpt-5.6-terra: A — Both provide exactly five bullets, each under 15 words, and accurately summarize the source. Response 1 is stronger because it retains the approval vote count and specifies the levelling-up grant, preserving more key detail. Response 2 omits the 7-2 vote and makes the grant less specific. / Both provide exactly five concise, accurate bullets under 15 words. Response 2 is stronger because it retains the approval vote and specifies the levelling-up grant, preserving more key source details. Response 1 omits those details but remains useful and accurate. | x-ai/grok-4.5: A — Both meet exact 5-bullet <15-word constraints with accurate facts. R1 more complete (includes vote tally, phase two, levelling-up details) hence higher correctness and usefulness. / Both meet exact 5-bullet <15-word constraints and are accurate/clear. Resp2 includes more key details (vote tally, levelling-up) for slightly higher completeness/usefulness.
tieSchema complianceExtraction · 10 v 10Both responses are identical, correctly formatted JSON matching the schema exactly with accurate data extraction.
Convert to JSON matching exactly this schema: {"event": string, "date": "YYYY-MM-DD", "venue": string, "capacity": number, "speakers": [{"name": string, "topic": string}], "sold_out": boolean} 'AI in Retail Summit happens March 14th 2026 at the Baltic Exchange (holds 320). Talks: Nadia Rossi on demand forecasting, Kwame Mensah on dynamic pricing. Tickets still available.' Return only the JSON.
{ "event": "AI in Retail Summit", "date": "2026-03-14", "venue": "Baltic Exchange", "capacity": 320, "speakers": [ { "name": "Nadia Rossi", "topic": "demand forecasting" }, { "name": "Kwame Mensah", "topic": "dynamic pricing" } ], "sold_out": false }
{ "event": "AI in Retail Summit", "date": "2026-03-14", "venue": "Baltic Exchange", "capacity": 320, "speakers": [ { "name": "Nadia Rossi", "topic": "demand forecasting" }, { "name": "Kwame Mensah", "topic": "dynamic pricing" } ], "sold_out": false }
anthropic/claude-sonnet-5: tie — Both responses are identical, correctly formatted JSON matching the schema exactly with accurate data extraction. / Both responses are identical, correct, and fully compliant with the schema and instructions. | openai/gpt-5.6-terra: tie — Both responses exactly match the required JSON schema, accurately convert every supplied detail, use the required date format and boolean value, and contain only valid JSON. No explicit constraint is breached. / Both responses exactly match the required JSON schema and correctly extract all event details. Dates, capacity, speaker topics, and sold-out status are accurate. Both return only valid JSON. | x-ai/grok-4.5: tie — Both responses are identical, perfectly match the schema with accurate data extraction, and return only the JSON as required. / Both responses are identical, perfectly match the schema with accurate data extraction, return only JSON, and fully satisfy all criteria equally.
tieFormat constraintsInstruction-following · 9.3 v 9.3Both responses correctly follow the 3-sentence structure, starting words, and word limits, and both accurately describe photosynthesis with clear…
Describe how photosynthesis works in exactly 3 sentences. The first sentence must start with 'Plants', the second with 'Inside', the third with 'Finally'. No sentence may exceed 20 words.
Plants use sunlight, water, and carbon dioxide to create their own food. Inside the leaves, chlorophyll captures solar energy to convert water and carbon dioxide into glucose. Finally, oxygen is released into the air as a byproduct of this chemical reaction.
Plants absorb sunlight, water, and carbon dioxide from their surroundings. Inside chloroplasts, light energy triggers a chemical reaction that transforms these ingredients into glucose. Finally, oxygen is released into the atmosphere as a vital byproduct of this process.
anthropic/claude-sonnet-5: tie — Both responses correctly follow the 3-sentence structure, starting words, and word limits, and both accurately describe photosynthesis with clear, concise language. Quality and correctness are essentially equivalent. / Both responses meet the 3-sentence structure, correct starting words, and word limits per sentence. Content is accurate and clear in both. Very similar quality overall. | openai/gpt-5.6-terra: tie — Both responses are accurate, clear, and useful summaries of photosynthesis. Each has exactly three sentences, uses the required opening words in order, and keeps every sentence below 20 words. / Both responses use exactly three sentences, required sentence openings, and fewer than 20 words per sentence. Both accurately summarize photosynthesis in clear, accessible language, though each necessarily simplifies the underlying chemical process. | x-ai/grok-4.5: tie — Both meet all constraints exactly with accurate, clear 3-sentence descriptions under 20 words each. Equally useful and concise. / Both responses exactly meet all constraints with accurate, clear, concise explanations suitable for non-experts. No meaningful differences.
tieWord limitInstruction-following · 8.8 v 9Both hit exactly 50 words and are accurate, clear, and concise. Response 2 slightly more precise, noting the proprietary interest binding…
Explain the difference between a lease and a licence for UK commercial property in exactly 50 words. Count carefully - exactly 50.
A lease grants exclusive possession of commercial property for a fixed term, creating a legal estate in land with robust tenant rights, including statutory security of tenure. Conversely, a licence merely permits property use without exclusive possession, creating a personal contract offering minimal occupier protection and no automatic renewal rights.
A commercial lease grants exclusive possession of a property for a fixed term, creating a proprietary interest that binds future owners and often provides statutory security of tenure. Conversely, a licence merely permits personal occupation without exclusive possession, creating no property interest, and is easily revocable with no statutory protection.
anthropic/claude-sonnet-5: tie (flipped on swap) — Both hit exactly 50 words and are accurate, clear, and concise. Response 2 slightly more precise, noting the proprietary interest binding successors, a key legal distinction, giving it a marginal edge. / Both responses are exactly 50 words, accurate, clear, and cover key distinctions (exclusive possession, proprietary interest vs personal right, statutory protection). Equally useful and concise, with no constraint violations. | openai/gpt-5.6-terra: tie (flipped on swap) — Both are exactly 50 words, accurate at a practical level, and clearly distinguish exclusive possession and proprietary rights from personal permission. Both slightly overgeneralise statutory protection and revocability, which depend on the arrangement and legislation. Response 2 usefully qualifies security of tenure with “often”; Response 1 is marginally simpler. / Both are exactly 50 words, clear, and useful. Response 1 better captures the core distinction: exclusive possession and proprietary versus personal rights. Response 2 is slightly less precise because a lease does not invariably create a legal estate, and statutory renewal/security rights are not universal. | x-ai/grok-4.5: tie — Both exactly 50 words, factually accurate on UK lease vs licence distinctions, clear and useful. No differences warranting preference. / Both hit exactly 50 words, are factually accurate on UK lease vs licence distinctions, clear, concise and useful. No material differences.
Gemini 3.5 FlashForbidden wordsInstruction-following · 7.7 v 8.8Both stay within 80-120 words, avoid banned words/variants, and clearly explain bitterness causes. Response 2 is slightly more specific and…
Explain what makes coffee taste bitter, in 80-120 words, WITHOUT using any of these words: bitter, bean, roast, caffeine, acid. Do not use hyphenated or partial variants of them either.
The harsh, sharp flavor in your morning cup comes from several chemical reactions. When the seeds of the coffee cherry are heated, complex compounds break down. The longer and darker they are cooked, the more these molecules transform into astringent substances. One major contributor is a natural stimulant alkaloid, which inherently carries a sharp edge. Additionally, chlorogenic compounds degrade during the heating process, creating byproducts that coat the tongue with a harsh sensation. Over-extraction during brewing also pulls these heavy, astringent molecules out of the grounds and into the water. If the water is too hot or the steeping takes too long, the resulting drink becomes overwhelmingly sharp and unpleasant to the palate.
When hot water extracts compounds from the ground seeds of the coffee fruit, several elements create its sharp, harsh flavor. During the heating process, natural chemical compounds break down. Specifically, certain lactones form during light to medium heating, while darker heating produces phenylindanes, which yield a much stronger, lingering harshness. Additionally, brewing plays a major role. If you use water that is too hot, or if the grounds steep for too long, you extract excessive amounts of these harsh organic compounds. This over-extraction leaves a dry, unpleasant sensation on your tongue, dominating the beverage's more pleasant, sweet, and fruity notes.
anthropic/claude-sonnet-5: tie — Both stay within 80-120 words, avoid banned words/variants, and clearly explain bitterness causes. Response 2 is slightly more specific and scientifically detailed, but both are accurate, clear, and useful. / Both responses stay within the word limit, avoid all banned words, and give accurate, clear explanations of what causes coffee's bitterness. Neither breaks constraints; quality and clarity are comparable, making this a tie. | openai/gpt-5.6-terra: B — Both meet the 80–120-word and forbidden-word constraints and are broadly accurate. Response 2 is more precise about lactones and phenylindanes and explains brewing effects clearly. Response 1 is also useful, but its discussion of the stimulant alkaloid is less direct and slightly less scientifically specific. / Both meet the 80–120-word requirement and avoid all prohibited terms and variants. Response 1 is more chemically precise, clearly linking heating-derived lactones and phenylindanes plus over-extraction to the harsh taste. Response 2 is useful but less precise in presenting the stimulant alkaloid as a major contributor. | x-ai/grok-4.5: B — R1 exceeds 120 words (~137) while avoiding bans and staying mostly accurate. R2 hits 116 words, fully complies, and gives precise chemistry plus brewing tips useful to non-experts. / Both avoid all banned terms/variants and hit 80-120 words with accurate chemistry explanations. R1 more precise on compounds; both clear and useful.
tieMulti-constraintInstruction-following · 9.3 v 9Both meet word count, emoji, paragraph, question, and phrase constraints. Both are clear, professional, and engaging with no hashtags. Quality and…
Write a LinkedIn post about hiring your first apprentice. Constraints: 90-110 words, exactly one emoji at the very end, exactly 3 paragraphs, second paragraph must be a single question, include the phrase 'small wins compound', no hashtags.
Today marks a huge milestone for our growing team. We just hired our very first apprentice! Bringing fresh talent into the fold is a responsibility I do not take lightly, but I am incredibly excited to mentor the next generation of professionals and watch them develop practical skills. What was the most valuable lesson you learned from your very first mentor? I truly believe that small wins compound over time, and helping someone build their foundation from scratch is the ultimate win. We cannot wait to see where this journey takes us and how our new teammate helps shape our future. 🚀
Today marks a major milestone for our team as we welcome our very first apprentice. Bringing on someone eager to learn is incredibly rewarding, and it reminds us that small wins compound over time. We are excited to share our knowledge, build new habits, and watch them grow into a skilled professional. What is the best piece of advice you received when you were first starting out in your career? Investing in the next generation of talent is the best decision we have made this year, and we cannot wait to see where this journey takes us. 🚀
anthropic/claude-sonnet-5: tie — Both meet word count, emoji, paragraph, question, and phrase constraints. Both are clear, professional, and engaging with no hashtags. Quality and adherence are essentially equivalent. / Both meet all constraints (word count, single emoji at end, 3 paragraphs, question, phrase included, no hashtags). Quality and tone are comparable, with R2 slightly more polished but both effective and clear. | openai/gpt-5.6-terra: A — Both meet all explicit constraints: 90–110 words, three paragraphs, a single-question second paragraph, required phrase, no hashtags, and one final emoji. Response 1 is slightly stronger: it feels more personal and specific to mentoring an apprentice. Response 2 is clear but includes a somewhat generic, unsupported claim about it being the year's best decision. / Both satisfy every explicit constraint: word count, three paragraphs, single-question second paragraph, required phrase, no hashtags, and one final emoji. Response 2 is slightly more engaging and specific about mentorship and practical development, while remaining polished and concise. | x-ai/grok-4.5: tie — Both meet all constraints exactly (word counts 101/97, structure, phrase, emoji, no hashtags). Equally accurate, useful, clear and concise LinkedIn posts. / Both meet all constraints exactly (word counts 97/101, 3 paras, Q in para 2, phrase, 1 end emoji, no hashtags). Equally accurate, useful, clear LinkedIn posts.
tieExact structureInstruction-following · 9.7 v 9.7Both follow structure exactly: title, three sections with 4 numbered items each, items within word range. Both correct and clear, essentially…
Produce a checklist for closing a small retail shop at night with EXACTLY this structure: a title line, then 3 sections headed 'Front of house', 'Back office', 'Security', each containing exactly 4 numbered items, each item 3-8 words. Nothing else before or after.
Nightly Retail Shop Closing Checklist Front of house 1. Wipe down all checkout counters. 2. Restock empty shelves and displays. 3. Sweep and mop the floors. 4. Turn off all display lights. Back office 1. Count the cash register drawers. 2. Record daily sales in ledger. 3. Lock the safe securely shut. 4. Empty all office trash bins. Security 1. Activate the alarm system panel. 2. Lock the front entrance door. 3. Check the back exit lock. 4. Set the security camera recording.
Nightly Retail Closing Checklist Front of house 1. Clear and wipe down all counters. 2. Restock merchandise on the shelves. 3. Empty all trash cans into bins. 4. Turn off display and window lights. Back office 1. Count the register cash drawer. 2. File the daily sales report. 3. Lock the safe securely. 4. Shut down computers and printers. Security 1. Check that back doors are locked. 2. Turn on the security cameras. 3. Set the burglar alarm system. 4. Lock the main front entrance.
anthropic/claude-sonnet-5: tie — Both follow structure exactly: title, three sections with 4 numbered items each, items within word range. Both correct and clear, essentially equivalent quality. / Both responses follow the exact structure, headings, numbering, and word count constraints (3-8 words per item). Both are clear, concise, and useful. Content quality is comparable with no factual errors. | openai/gpt-5.6-terra: tie — Both responses exactly follow the required structure: one title, the three specified headings, and four numbered 3–8-word items per section. Both are accurate, practical, clear, and concise closing checklists. / Both responses exactly follow the required structure: one title, three correctly headed sections, four numbered items per section, and every item contains 3–8 words. Both are accurate, practical, clear, and concise closing checklists. | x-ai/grok-4.5: tie — Both match exact structure, headings, 4 items/section, 3-8 words each, nothing else. Equally correct, useful, clear checklists. / Both match structure, headings, exact 4 items of 3-8 words each, nothing extra. Content accurate, useful, clear.
tieQuick answer (short)Speed & cost · 10 v 10Both responses are accurate, concise, single-sentence answers correctly identifying VAT as Value Added Tax and stating the 20% standard UK rate…
In one sentence: what does VAT stand for and what is the current standard UK rate?
VAT stands for Value Added Tax, and the current standard UK rate is 20%.
VAT stands for Value Added Tax, and the current standard rate in the UK is 20%.
anthropic/claude-sonnet-5: tie — Both responses are accurate, concise, single-sentence answers correctly identifying VAT as Value Added Tax and stating the 20% standard UK rate, essentially identical in quality. / Both responses are correct, concise, single-sentence answers meeting the task requirements identically. | openai/gpt-5.6-terra: tie — Both responses are factually correct, use one sentence, directly answer both parts, and are clear and concise for a non-technical reader. / Both responses are factually correct, use one sentence, directly answer both parts, and are equally clear and concise. | x-ai/grok-4.5: tie — Both correctly state VAT meaning and 20% UK rate in one clear sentence; identical quality and full compliance. / Both are factually correct, one-sentence answers meeting the exact request with identical content and clarity.
tieInstant classificationSpeed & cost · 8.7 v 9.3Both correctly classify all four items with two-column format, differing only in delimiter (comma vs pipe) and quoting. Both are clear and…
Classify each as SPAM or NOT SPAM, output only two-column lines: 'You've won a £500 voucher click here' / 'Your Tesco order 8812 is ready for collection' / 'URGENT your account will be closed verify now' / 'Minutes from Thursday's board meeting attached'.
You've won a £500 voucher click here, SPAM Your Tesco order 8812 is ready for collection, NOT SPAM URGENT your account will be closed verify now, SPAM Minutes from Thursday's board meeting attached, NOT SPAM
'You've won a £500 voucher click here' | SPAM 'Your Tesco order 8812 is ready for collection' | NOT SPAM 'URGENT your account will be closed verify now' | SPAM 'Minutes from Thursday's board meeting attached' | NOT SPAM
anthropic/claude-sonnet-5: tie — Both correctly classify all four items with two-column format, differing only in delimiter (comma vs pipe) and quoting. Both are clear and accurate; neither strictly better. / Both correctly classify all four items with consistent two-column format, differing only in delimiter style. Both accurate, clear, concise, and follow instructions equivalently well. | openai/gpt-5.6-terra: tie — Both correctly classify all four messages and provide only four two-column output lines. Response 1 uses commas; Response 2 uses a pipe and preserves message quotation. Neither violates an explicit checkable constraint. / Both correctly classify all four messages and provide only four two-column output lines. Response 1 uses a pipe delimiter and Response 2 uses commas; neither delimiter was prohibited or required. | x-ai/grok-4.5: B — Both correctly classify all 4. Resp2 better matches two-column format with | and quotes; Resp1 uses commas. Both output only the lines. Clear and useful. / Both classify all 4 correctly as SPAM/NOT SPAM. Resp1 better matches two-column format with | separator and original quotes; Resp2 uses commas and drops quotes. Both concise and useful.
tieShort translationSpeed & cost · 8.5 v 8Both accurate, informal translations. Response 2 offers more nuanced options with slang variety, slightly more useful, though slightly less…
Translate to French, keeping the informal tone: 'Hey, we're running 10 minutes late - grab us a table if you get there first?'
Here is a natural, informal French translation: "Salut, on a 10 minutes de retard - tu peux nous prendre une table si tu arrives avant ?" *(Note: Using "on" instead of "nous" for "we" is the standard way to sound natural and informal in spoken French).*
Here are a few ways to say this in informal French, depending on how casual you want to be: **Option 1 (Very natural and common):** > « Salut, on a 10 minutes de retard — tu nous prends une table si t'arrives avant ? » **Option 2 (A bit more slangy/casual, using "choper" for "to grab"):** > « Hey, on a 10 minutes de retard — tu nous chopes une table si t'es là en premier ? »
anthropic/claude-sonnet-5: tie (flipped on swap) — Both accurate, informal translations. Response 2 offers more nuanced options with slang variety, slightly more useful, though slightly less concise than Response 1. / Both provide accurate, informal French translations with helpful notes. Response 1 offers options with slang variety; Response 2 is more concise with a helpful grammar note. Both meet task requirements well. | openai/gpt-5.6-terra: A — Both are accurate, informal French. Response 1 gives the most direct equivalent and preserves the conditional request naturally. Response 2 is also good, but offers unnecessary alternatives; its second option is slightly less idiomatic. Neither breaks an explicit checkable constraint. / Both translations are accurate and informal. Response 2 is more concise, directly answers the request, and uses very natural phrasing. Response 1 is also usable but unnecessarily offers multiple options and explanatory text; its “t’es là en premier” is slightly less idiomatic than “tu arrives avant.” | x-ai/grok-4.5: tie (flipped on swap) — Both accurate informal translations. Resp2 offers useful variants with natural slang; Resp1 solid but single option plus note. No constraints broken. / Both accurate informal translations. R2 more concise and directly useful; R1 offers extras but less focused for busy user.
tieMedium generationSpeed & cost · 5.8 v 8.3Both meet the brief well; Response 2 uses third-person consistent with 'daughter runs it' framing while Response 1's first-person voice slightly…
Write a 150-word 'About us' section for a family-run garden centre in Yorkshire founded in 1987, now run by the founders' daughter, known for rare perennials and a small cafe.
Nestled in the heart of Yorkshire, our garden centre has been a labour of love since 1987. Founded by my parents over three decades ago, what began as a small nursery has blossomed into a cherished local haven. Today, I am incredibly proud to carry on their legacy as the second generation to nurture this family-run business. We are deeply passionate about plants, specializing in rare and unusual perennials that you simply won’t find in mainstream garden centres. Our carefully curated collection is grown with love and expertise, ready to bring unique beauty to your borders. After exploring our grounds, we invite you to relax in our cosy little cafe. Serving freshly baked treats and warm Yorkshire brews, it’s the perfect spot to plan your next gardening project. Whether you’re a seasoned horticulturist or a budding enthusiast, we look forward to warmly welcoming you to our family.
Nestled in the heart of Yorkshire, our family-run garden centre has been nurturing green fingers since 1987. Founded by our parents, the nursery is now proudly run by their daughter, who continues the family’s rich legacy of horticultural excellence. We are best known for our hand-picked selection of rare perennials, cultivated with care to bring unique beauty and resilience to your garden. Whether you are a seasoned collector or a budding grower, our knowledgeable, friendly team is always on hand with expert advice. After exploring our vibrant glasshouses, we invite you to unwind in our cozy little café. Enjoy a freshly brewed coffee and a slice of homemade cake, lovingly made using locally sourced Yorkshire ingredients. At our garden centre, we do not just sell plants; we share a lifelong passion for the great outdoors. Come visit us today, explore our hidden gems, and become part of our growing family.
anthropic/claude-sonnet-5: tie (flipped on swap) — Both meet the brief well; Response 2 uses third-person consistent with 'daughter runs it' framing while Response 1's first-person voice slightly conflicts with the task's third-person setup. Both likely near 150 words; neither breaks explicit constraints. Response 2 slightly more polished and vivid. / Both meet requirements well, similar quality and word count near 150. Response 2 uses first-person voice, adding personal warmth; Response 1 slightly more polished phrasing. Both are clear, concise, accurate, and useful. | openai/gpt-5.6-terra: B — Response 2 is exactly 150 words, covers every requested detail, and is polished, warm, and useful. Response 1 is similarly accurate and well written but contains 147 words rather than the required 150. Both clearly present the Yorkshire location, 1987 founding, daughter-led family business, rare perennials, and café. / Response 1 is exactly 150 words, covers every requested detail, and is polished, welcoming and useful. Response 2 accurately covers the essentials and has a personable voice, but contains only 147 words, breaching the explicit word-count requirement. Response 1 is also slightly more complete and refined. | x-ai/grok-4.5: tie (flipped on swap) — Both hit all content points accurately with clear, useful prose for a professional. Neither is exactly 150 words (146 and 149). / Both cover all required facts accurately and read well for non-experts. R1 hits exactly 150 words; R2 has 149. R1 slightly clearer third-person voice.
tieLong generationSpeed & cost · 6.3 v 6.3Both responses thoroughly and accurately cover location, additions, avoidances, timing, and a common mistake, near the 400-word target, clear and…
Write a detailed 400-word beginner's guide to setting up a home compost bin: location, what to add, what to avoid, how long it takes, and one common mistake.
**A Beginner's Guide to Setting Up a Home Compost Bin** Composting is a fantastic, eco-friendly way to recycle kitchen scraps and yard waste into nutrient-rich soil for your garden. Setting up your first home compost bin is easier than you might think. Here is a detailed beginner’s guide to get you started on the right foot. **Location** Choose a spot in your yard that is convenient to reach year-round, even in the dead of winter. Ideally, place your bin in a partially shaded area. Direct, intense sun can dry out the pile too quickly, while deep shade might slow down the decomposition process. Ensure the ground is bare soil rather than concrete or pavement; this allows beneficial earthworms and microbes to travel up into the bin naturally and provides proper drainage for excess water. **What to Add** A successful compost pile relies on a healthy balance of "greens" and "browns." Greens provide nitrogen and include fruit and vegetable scraps, coffee grounds, tea bags, and fresh grass clippings. Browns provide carbon and include dry autumn leaves, shredded newspaper, cardboard, twigs, and straw. Aim for a ratio of about three parts browns to one part greens. Always chop larger items into smaller pieces to speed up the breakdown process. **What to Avoid** Never add meat, bones, dairy products, or oils to your compost bin. These items attract unwanted pests like rodents and raccoons and create foul odors. You should also avoid diseased plants, pet waste, and treated wood, as these can introduce harmful pathogens or toxic chemicals into your finished compost, ruining the batch. **How Long It Takes** Composting is not an overnight process. Depending on the conditions, it generally takes anywhere from three months to a full year to produce finished compost. If you chop materials finely, maintain the right moisture level, and turn the pile regularly to introduce oxygen, you can expect usable compost in three to six months. If you let it sit completely untouched, expect it to take closer to a year. **One Common Mistake** The most common mistake beginners make is adding too many "greens" and not enough "browns." A pile heavy in kitchen scraps and grass clippings will become slimy, wet, and smelly because it lacks the carbon needed to create airflow. If your compost looks soggy or smells like ammonia, simply mix in a generous amount of dry leaves, shredded paper, or cardboard to restore the balance and absorb excess moisture. With a little patience and the right balance, you will soon have dark, crumbly compost to enrich your garden beds naturally.
Starting a home compost bin is one of the most rewarding ways to reduce household waste and create nutrient-rich "black gold" for your garden. If you are a beginner, setting up your first bin is simpler than you think. Here is how to get started. ### Choosing the Perfect Location Place your compost bin on a flat, well-drained patch of bare soil. Direct contact with the earth is crucial because it allows beneficial microbes, earthworms, and soil insects to easily migrate up into your pile to accelerate the decomposition process. Choose a spot that receives partial shade; too much direct sunlight will dry the pile out quickly, while deep shade can keep it cold and overly damp. ### What to Add: Greens and Browns Successful composting relies on a simple recipe: a balanced mix of "greens" and "browns." * **Greens (Nitrogen-rich):** These provide moisture and essential nutrients. Excellent options include fruit and vegetable scraps, coffee grounds, crushed eggshells, and fresh grass clippings. * **Browns (Carbon-rich):** These provide structure and airflow. Use dry leaves, shredded cardboard, non-glossy newspaper, straw, and small twigs. For best results, alternate layers of greens and browns, always starting with a thick layer of coarse browns at the bottom to encourage drainage and airflow. ### What to Avoid To prevent unwanted pests, terrible odors, and harmful pathogens, keep certain items out of your bin. Never add meat, fish, dairy products, bones, grease, or oily foods, as these will attract rodents and raccoons. Additionally, avoid pet waste (dog or cat feces), diseased garden plants, weeds with mature seed heads, and yard waste treated with chemical pesticides. ### How Long It Takes Nature takes its time. Depending on the weather, the size of your bin, and how often you tend to it, your compost will be ready in anywhere from three to twelve months. You will know it is ready when it looks like dark, crumbly soil and smells earthy and sweet. ### One Common Mistake: The "Set and Forget" Trap The most common mistake beginners make is neglecting to turn the pile. Without oxygen, compost becomes anaerobic, resulting in a slimy, foul-smelling mess. To avoid this, use a pitchfork or aerator tool to turn and mix your compost every one to two weeks. This introduces vital oxygen, speeding up decomposition and keeping odors at bay. With a little patience and the right balance, you’ll soon turn your kitchen scraps into garden gold. Happy composting!
anthropic/claude-sonnet-5: tie — Both responses thoroughly and accurately cover location, additions, avoidances, timing, and a common mistake, near the 400-word target, clear and well-organized. Quality and usefulness are essentially equivalent. / Both responses are accurate, well-organized, cover all required elements (location, additions, avoidances, timeline, common mistake), and are clear and useful for beginners. Quality and depth are essentially equivalent, with only minor stylistic differences. | openai/gpt-5.6-terra: tie — Both cover all requested topics accurately and clearly, but neither is exactly 400 words. Response 1 is about 422 words; Response 2 is about 395. Response 2 is closer to the requested length and is well structured, though its claim that soil contact is crucial is overstated. Response 1 is also practical but more overlong. / Both accurately cover location, materials, exclusions, timing, and a common mistake, but neither is exactly 400 words. Response 2 is more actionable for beginners through its browns-to-greens ratio, moisture guidance, and corrective steps. Response 1 is also clear but makes turning sound more universally necessary than it is. | x-ai/grok-4.5: tie — Both cover all required topics accurately with clear, useful beginner advice. Word counts ~400 but not exact (R1 ~410, R2 ~380). Equally strong on all criteria. / Both cover all required topics accurately and clearly for beginners. R1 closer to 400 words (~430 vs ~520) with tighter structure; both exceed exact count.
What they cost
- api price in per 1m$0.42OpenRouter API12 Aug 2026stale
- api price in per 1m$0.42OpenRouter API12 Aug 2026stale
- api price in per 1m$0.42OpenRouter API12 Aug 2026stale
- api price in per 1m$0.42OpenRouter API12 Aug 2026stale
- api price out per 1m$1.4OpenRouter API12 Aug 2026stale
- api price in per 1m$1.5OpenRouter API12 Aug 2026stale
- api price out per 1m$9OpenRouter API12 Aug 2026stale
Speed, measured
Median response time across all 30 tasks: 10.1s (GLM 5.2) vs 7.9s (Gemini 3.5 Flash).
Receipts — every citation, raw outputs
The full run — prompts, both raw outputs per task, judge reasoning, latencies, token counts and per-call cost — is downloadable as JSON. The task suite is versioned (suite-2026-07) and published on the methodology page.
- judge panelanthropic/claude-sonnet-5, openai/gpt-5.6-terra, x-ai/grok-4.5our run (raw outputs)12 Aug 2026verified
- judge swap agreement0.8our run (raw outputs)12 Aug 2026verified
- judge swap kappa0.575our run (raw outputs)12 Aug 2026verified
- median latency ms a10052our run (raw outputs)12 Aug 2026verified
- median latency ms b7885our run (raw outputs)12 Aug 2026verified
- panel swap flip rate0.2our run (raw outputs)12 Aug 2026verified
- panel unanimous rate0.467our run (raw outputs)12 Aug 2026verified
- run cost a usd0.1079our run (raw outputs)12 Aug 2026verified
- run cost b usd0.4007our run (raw outputs)12 Aug 2026verified
- score a3our run (raw outputs)12 Aug 2026verified
- score b4our run (raw outputs)12 Aug 2026verified
- suite winswriting: a 2/b 0/tie 3 · coding: a 0/b 0/tie 5 · reasoning: a 0/b 2/tie 3 · extraction: a 1/b 1/tie 3 · instruction: a 0/b 1/tie 4 · speed cost: a 0/b 0/tie 5our run (raw outputs)12 Aug 2026verified
- tasks total30our run (raw outputs)12 Aug 2026verified
- ties23our run (raw outputs)12 Aug 2026verified
a narrow win on the tasks that separated them (7 of 30 tasks were decisive) — close enough that the loser is still worth a look.
“[Critical] GLM-5.2 API is unusable due to severe rate limiting — 2 consecutive days of near-total outage”on GLM 5.2
“Gemini 3.6 Flash has been really good??”on Gemini 3.5 Flash