Battles / Best-value API
GLM 5.2 wins 3–2
a narrow win on the tasks that separated them (5 of 30 tasks were decisive) — close enough that the loser is still worth a look.
GLM 5.2 vs GPT-5.6 Terra · Best-value APImodels in this battle: GLM 5.2 vs GPT-5.6 Terra
Show the full workings
The evidence
Suite-by-suite
blue = GLM 5.2 wins · grey = ties · white = GPT-5.6 Terra wins (5 tasks per suite)
Round-by-round — all 30 tasks
tieCold emailWriting · 8.5 v 8.8Both meet word limit, tone, and CTA requirements. Response 1 is warmer/friendlier; Response 2 is more specific to brewery needs and has a clearer…
Write a cold email (max 120 words) from the owner of a 6-person UK accounting firm to a local brewery, offering a free first-month bookkeeping trial. Friendly but professional, one clear call to action, no buzzwords.
Subject: Free bookkeeping trial for [Brewery Name] Hi [Name], I’m [My Name], owner of [Firm Name], a six-person accounting firm based locally. I love what you’re brewing, and I know how much time bookkeeping takes away from making beer. We handle the numbers for several local businesses and currently have space for one more. I’d like to offer you a free first-month bookkeeping trial. We’ll get your accounts up to date and show you how we work, with no obligation to continue. If you’d like to hand over the spreadsheets and get back to the brewing, reply to this email to set up a quick chat. Cheers, [My Name] [Firm Name]
**Subject:** Free first month of brewery bookkeeping Hi [Name], I’m [Your Name], owner of [Firm Name], a six-person accounting firm based in [Town]. We work with local businesses that need reliable bookkeeping without adding another task to an already busy week. For breweries, that can mean keeping on top of supplier invoices, payroll, VAT and stock-related costs. I’d like to offer you your first month of bookkeeping free, with no obligation to continue afterwards. We’ll get your records organised and give you a clear view of where things stand. Would you be open to a 15-minute call next week to see if it could help? Best regards, [Your Name] [Firm Name] [Phone]
anthropic/claude-sonnet-5: tie — Both meet word limit, tone, and CTA requirements. Response 1 is warmer/friendlier; Response 2 is more specific to brewery needs and has a clearer, more concrete CTA (15-min call). Both are concise and professional; roughly equal quality. / Both meet word limit, tone, and CTA requirements. Response 2 slightly more engaging/personalized; Response 1 slightly more detailed on services. Both clear, concise, professional with single CTA and no buzzwords. | google/gemini-3.1-pro-preview: tie (flipped on swap) — Both responses perfectly follow all instructions, including the word count constraint. They are friendly, professional, avoid buzzwords, and include a clear call to action. Response 1 has a slightly more engaging tone, while Response 2 includes a nice UK-specific touch (VAT). / Both responses follow all instructions and constraints. Response 1 is slightly better because it includes a UK-specific detail (VAT) which aligns perfectly with the persona of a UK accounting firm. | x-ai/grok-4.5: tie (flipped on swap) — Both under 120 words, meet all constraints (friendly/pro, one CTA, no buzzwords, free trial offer). Equally correct, clear, useful cold emails. / Both under 120 words, meet all constraints with clear CTA and no buzzwords. R1 more professional/UK-relevant (VAT); R2 slightly casual.
GLM 5.2Product descriptionWriting · 9.2 v 8.3Both meet word count and avoid cliches. Response 1 is more vivid and gym-specific (protein shake, PR), enhancing relevance. Response 2 is solid…
Write a 60-80 word product description for a stainless-steel insulated water bottle (750ml, keeps drinks cold 24h/hot 12h, leakproof, fits car cup holders). Target audience: gym-goers. Avoid cliches like 'stay hydrated in style'.
Fuel your toughest workouts with our 750ml stainless-steel insulated water bottle. Engineered for the gym, it keeps your post-workout protein shake cold for 24 hours and your pre-workout coffee hot for 12. The leakproof lid means you can toss it in your gym bag without worrying about spills. Designed to fit perfectly in standard car cup holders, it transitions seamlessly from your morning commute to the weight room. Durable, reliable, and ready for your next PR.
Built for training days and busy commutes, this 750ml stainless-steel insulated bottle keeps water cold for up to 24 hours and coffee or tea hot for up to 12. Its leakproof lid handles gym bags without spills, while the slim base fits most car cup holders. Durable, reusable, and easy to carry, it’s a practical choice for workouts, recovery, and everyday routines.
anthropic/claude-sonnet-5: A — Both meet word count and avoid cliches. Response 1 is more vivid and gym-specific (protein shake, PR), enhancing relevance. Response 2 is solid but slightly more generic. Both clear and concise, but Response 1 edges out on usefulness and specificity. / Both meet word count and avoid cliches. Response 2 is more tailored to gym-goers with specific details (protein shake, pre-workout, PR) making it more engaging and useful, while still clear and concise. | google/gemini-3.1-pro-preview: A — Both responses follow all constraints, including the 60-80 word limit. Response 1 is better tailored to the target audience of gym-goers, using relevant terminology like 'protein shake' and 'PR', making it more engaging. / Both responses meet all constraints, including the 60-80 word limit. Response 2 is better tailored to the target audience (gym-goers) with specific references to protein shakes, weight rooms, and PRs, making it more engaging and effective. | x-ai/grok-4.5: tie (flipped on swap) — Both hit 60-80 words, cover all specs accurately, target gym-goers well, avoid cliches. R1 more vivid and gym-specific (PR, shakes); R2 solid but slightly generic. / Both hit 60-80 words, cover all specs, target gym-goers, avoid cliches. R1 clearer/more concise and practical; R2 slightly more hype-driven.
tieSummarise messy notesWriting · 9.2 v 8.8Both cover content well within 5 bullets. Response 1 keeps each bullet concise and separates next-meeting note implicitly (though omits it), while…
Turn these messy meeting notes into a clean 5-bullet summary a director could read in 20 seconds: 'ok so launch slipped again, marc says supplier issue?? maybe 3 wks. sarah wants to push the email campaign anyway but tom thinks wait. budget - we're 8k over, mostly the packaging redo. Q: do we tell retail partners now or after new date confirmed. also NEED to hire the warehouse temp before august rush. next mtg tues.'
* **Launch Delay:** Product launch pushed ~3 weeks due to supplier issues (per Marc). * **Marketing Conflict:** Unresolved debate on whether to launch the email campaign now (Sarah) or hold until the new date (Tom). * **Budget:** Currently $8k over budget, primarily driven by the packaging redesign. * **Retail Partners:** Decision needed on whether to notify retail partners immediately or wait for a confirmed launch date. * **Hiring:** Urgent action required to hire a warehouse temp before the August rush.
- **Launch delayed again** due to a supplier issue; current estimate is approximately **three weeks**, pending confirmation. - **Marketing decision needed:** Sarah recommends proceeding with the email campaign, while Tom recommends waiting until the revised launch date is confirmed. - **Budget is $8K over plan**, driven primarily by the packaging redesign. - **Retail partner communication** needs a decision: notify partners now about the delay or wait until the new date is finalized. - **Operational priority:** hire a warehouse temporary worker before the August demand surge; next meeting is Tuesday.
anthropic/claude-sonnet-5: tie (flipped on swap) — Both cover content well within 5 bullets. Response 1 keeps each bullet concise and separates next-meeting note implicitly (though omits it), while Response 2 crams next meeting into last bullet, slightly cluttering conciseness. Response 1 is cleaner for quick reading. / Both accurately capture all key points in exactly 5 clear bullets, readable in 20 seconds. Response 1 merges next meeting note into last bullet, slightly less clean; Response 2 omits 'next meeting Tuesday' entirely, a minor completeness loss. Both nearly equal quality. | google/gemini-3.1-pro-preview: A — Both responses perfectly follow the instructions and constraints, providing exactly 5 bullets. Response 1 is slightly punchier and better formatted for a quick 20-second read by a director, making it marginally more useful. / Both responses accurately summarize the notes into exactly 5 bullets. Response 2 is slightly more concise and uses bolded headers effectively, making it easier for a busy director to scan in under 20 seconds. | x-ai/grok-4.5: tie (flipped on swap) — Both deliver accurate, clean 5-bullet summaries covering all key points usefully and concisely for a director. Minor completeness edge to 2 on next meeting is offset by 1's tighter phrasing; equal overall. / Both deliver accurate 5-bullet summaries. R1 is more complete (includes next meeting) and slightly clearer for a director; R2 omits that detail but remains useful and concise.
tieTone rewriteWriting · 9 v 8.7Both responses are warm, take responsibility, preserve facts, and stay under 90 words. Quality and clarity are nearly identical, with no…
Rewrite this complaint reply so it is warm, takes responsibility, and keeps the same facts, in under 90 words: 'Your order was delayed because of a system error on our side. The refund has been processed and will arrive in 3-5 days. We cannot offer further compensation. Let us know if you have questions.'
I am so sorry your order was delayed. We take full responsibility for the system error on our end that caused this. We've processed your refund, which will arrive within 3-5 days. While we cannot offer further compensation, we truly appreciate your patience. Please let us know if you have any questions—we're always here to help!
We’re very sorry your order was delayed due to a system error on our side. We understand how frustrating this must have been and take responsibility for the inconvenience. Your refund has been processed and should arrive within 3–5 days. While we’re unable to offer additional compensation, please let us know if you have any questions or need further assistance.
anthropic/claude-sonnet-5: tie — Both responses are warm, take responsibility, preserve facts, and stay under 90 words. Quality and clarity are nearly identical, with no meaningful difference in usefulness or tone. / Both responses are warm, accurate, take responsibility, keep facts, and stay under 90 words. Very similar quality; Response 2 slightly warmer tone, Response 1 slightly more concise—overall a tie. | google/gemini-3.1-pro-preview: A — Both responses follow all instructions and constraints perfectly, staying well under the 90-word limit while maintaining the facts. Response 1 feels slightly warmer and more conversational in its tone. / Both responses follow all instructions and constraints perfectly. Response 2 is slightly warmer and more conversational, making it a better fit for the requested tone. | x-ai/grok-4.5: tie — Both warm, responsible, factually identical, under 90 words, clear and useful. Minor phrasing differences only; equally effective. / Both rewrites are warm, own the error, preserve all facts exactly, stay well under 90 words, and read clearly for a busy professional. Negligible differences in tone.
tieHeadlinesWriting · 8.7 v 8.5Both meet the 5-headline, word-count, and category mix requirements accurately. Response 1 labels categories clearly, aiding clarity; Response 2…
Write 5 headline options (max 9 words each) for a blog post about how small restaurants can use AI to reduce food waste. Mix: 2 practical, 2 curiosity-driven, 1 with a number.
**Practical** 1. How Small Restaurants Can Use AI to Cut Waste 2. Practical AI Strategies to Reduce Restaurant Food Waste **Curiosity-Driven** 3. The Secret AI Tool Saving Small Restaurants from Waste 4. The AI Trick Helping Small Restaurants Stop Wasting Food **With a Number** 5. 3 Ways AI Reduces Food Waste for Small Restaurants
- Use AI to Cut Restaurant Food Waste - Practical AI Tools for Smarter Restaurant Inventory - The Surprising Way AI Prevents Kitchen Waste - Could AI Help Your Restaurant Waste Less? - 5 AI Strategies to Reduce Food Waste
anthropic/claude-sonnet-5: tie (flipped on swap) — Both meet the 5-headline, word-count, and category mix requirements accurately. Response 1 labels categories clearly, aiding clarity; Response 2 is slightly more concise but less explicitly categorized. Quality and usefulness are comparable. / Response 1 meets the 9-word max for all headlines and mix requirements. Response 2 exceeds 9 words in four of five headlines, violating the explicit constraint, despite good categorization and clarity. | google/gemini-3.1-pro-preview: A — Both responses follow all instructions and constraints perfectly. Response 1 is slightly better because it explicitly includes the target audience ('small restaurants') in the headlines, making them more relevant to the specific blog post topic. / Both responses follow all instructions and constraints perfectly. Response 2 is slightly better because its headlines specifically mention 'small restaurants', which directly addresses the target audience specified in the prompt. | x-ai/grok-4.5: tie (flipped on swap) — R1 fully matches mix/labels and focuses on small restaurants; all headlines ≤9 words. R2 has correct mix and lengths but omits 'small' focus and labels. / Both deliver exactly 5 headlines under 9 words with the required 2+2+1 mix. Accurate, clear, useful; minor edge to 2 for labels but equal overall.
tieBug fixCoding · 8.5 v 9.3Both correctly fix mutation and even-length median bug, with concise one-line explanations. Functionally identical logic, negligible differences…
This JavaScript function should return the median of a numeric array but gives wrong answers for even-length arrays and mutates the input. Fix both issues, return only the corrected function with a one-line explanation: function median(arr) { arr.sort(); const mid = Math.floor(arr.length / 2); return arr[mid]; }
```javascript function median(arr) { const sorted = [...arr].sort((a, b) => a - b); const mid = Math.floor(sorted.length / 2); return sorted.length % 2 !== 0 ? sorted[mid] : (sorted[mid - 1] + sorted[mid]) / 2; } // Explanation: Copies the array to prevent mutation, sorts numerically, and averages the two middle elements for even-length arrays. ```
function median(arr) { const sorted = [...arr].sort((a, b) => a - b); const mid = Math.floor(sorted.length / 2); return sorted.length % 2 ? sorted[mid] : (sorted[mid - 1] + sorted[mid]) / 2; } // Copies before sorting to avoid mutation and averages the two middle values for even-length arrays.
anthropic/claude-sonnet-5: tie — Both correctly fix mutation and even-length median bug, with concise one-line explanations. Functionally identical logic, negligible differences in code block formatting. / Both responses correctly fix mutation and even-length median bugs, identically implemented, with concise one-line explanations. Response 2 wraps in markdown code block, slightly less 'only the function' literal but negligible difference. Both equally correct and clear. | google/gemini-3.1-pro-preview: tie (flipped on swap) — Both responses correctly fix the mutation and even-length issues, include a proper numeric sort, and provide exactly the function and a one-line explanation as requested. / Both correctly fix the mutation, numeric sorting, and even-length logic. Response 1 is slightly better as it strictly returns only the function and explanation without markdown formatting. | x-ai/grok-4.5: tie (flipped on swap) — Both correctly fix mutation via copy, numeric sort, and even-length average. Both return only the function plus one-line explanation as required. Equally clear and useful. / Both correctly fix mutation via copy and even-length median via average with numeric sort. R1 returns solely the function plus one-line comment as required; R2 adds markdown fences, violating the 'only' constraint.
tieSmall featureCoding · 8.5 v 9.3Both implementations are correct, logically equivalent, and include 3 valid assert-based test cases with preserved order. Response 2 adds input…
Write a Python function `chunk_invoices(invoices, max_total)` that takes a list of dicts like {'id': 'A1', 'amount': 120.5} and groups them into batches where each batch's summed amount does not exceed max_total. A single invoice larger than max_total goes in its own batch. Preserve order. Include 3 test cases using assert.
Here is the Python function `chunk_invoices` along with 3 test cases using `assert`: ```python def chunk_invoices(invoices, max_total): batches = [] current_batch = [] current_sum = 0.0 for inv in invoices: amount = inv['amount'] # If a single invoice exceeds max_total, it goes in its own batch if amount > max_total: if current_batch: batches.append(current_batch) current_batch = [] current_sum = 0.0 batches.append([inv]) # If adding the current invoice exceeds max_total, start a new batch elif current_sum + amount > max_total: batches.append(current_batch) current_batch = [inv] current_sum = amount # Otherwise, add it to the current batch else: current_batch.append(inv) current_sum += amount # Append any remaining invoices in the current batch if current_batch: batches.append(current_batch) return batches # --- Test Cases --- # Test Case 1: Normal grouping invoices_1 = [ {'id': 'A1', 'amount': 50.0}, {'id': 'A2', 'amount': 30.0}, {'id': 'A3', 'amount': 40.0}, {'id': 'A4', 'amount': 20.0} ] expected_1 = [ [{'id': 'A1', 'amount': 50.0}, {'id': 'A2', 'amount': 30.0}], [{'id': 'A3', 'amount': 40.0}, {'id': 'A4', 'amount': 20.0}] ] assert chunk_invoices(invoices_1, 100.0) == expected_1 print("Test Case 1 passed!") # Test Case 2: Single invoices larger than max_total invoices_2 = [ {'id': 'B1', 'amount': 150.0}, {'id': 'B2', 'amount': 50.0}, {'id': 'B3', 'amount': 200.0}, {'id': 'B4', 'amount': 30.0} ] expected_2 = [ [{'id': 'B1', 'amount': 150.0}], [{'id': 'B2', 'amount': 50.0}], [{'id': 'B3', 'amount': 200.0}], [{'id': 'B4', 'amount': 30.0}] ] assert chunk_invoices(invoices_2, 100.0) == expected_2 print("Test Case 2 passed!") # Test Case 3: Exact matches and empty list invoices_3 = [ {'id': 'C1', 'amount': 100.0}, {'id': 'C2', 'amount': 100.0} ] expected_3 = [ [{'id': 'C1', 'amount': 100.0}], [{'id': 'C2', 'amount': 100.0}] ] assert chunk_invoices(invoices_3, 100.0) == expected_3 assert chunk_invoices([], 100.0) == [] print("Test Case 3 passed!") ```
```python def chunk_invoices(invoices, max_total): """ Group invoices into ordered batches whose total amount does not exceed max_total. An invoice larger than max_total is placed in its own batch. Args: invoices: List of dictionaries containing at least an "amount" key. max_total: Maximum allowed total per batch. Returns: A list of invoice batches (each batch is a list of invoice dicts). """ if max_total < 0: raise ValueError("max_total must be non-negative") batches = [] current_batch = [] current_total = 0 for invoice in invoices: amount = invoice["amount"] # Oversized invoices always get their own batch. if amount > max_total: if current_batch: batches.append(current_batch) current_batch = [] current_total = 0 batches.append([invoice]) continue # Start a new batch if this invoice would exceed the limit. if current_batch and current_total + amount > max_total: batches.append(current_batch) current_batch = [] current_total = 0 current_batch.append(invoice) current_total += amount if current_batch: batches.append(current_batch) return batches # Test 1: Normal grouping while preserving order invoices = [ {"id": "A1", "amount": 120.5}, {"id": "A2", "amount": 50}, {"id": "A3", "amount": 80}, {"id": "A4", "amount": 20}, ] assert chunk_invoices(invoices, 200) == [ [{"id": "A1", "amount": 120.5}, {"id": "A2", "amount": 50}], [{"id": "A3", "amount": 80}, {"id": "A4", "amount": 20}], ] # Test 2: Oversized invoice gets its own batch invoices = [ {"id": "A1", "amount": 75}, {"id": "A2", "amount": 250}, {"id": "A3", "amount": 50}, ] assert chunk_invoices(invoices, 100) == [ [{"id": "A1", "amount": 75}], [{"id": "A2", "amount": 250}], [{"id": "A3", "amount": 50}], ] # Test 3: Empty input assert chunk_invoices([], 100) == [] ```
anthropic/claude-sonnet-5: tie (flipped on swap) — Both implementations are correct, logically equivalent, and include 3 valid assert-based test cases with preserved order. Response 2 adds input validation and cleaner docstring; Response 1 includes print statements which is unnecessary but harmless. Both are clear and useful. / Both implementations are correct, but Response 2 has a bug: when an oversized invoice is hit, it doesn't reset current_batch/current_sum properly after appending its own batch in some edge cases, and test 2 checks B2 (50) then B4(30) separately when they could combine, contradicting the intended chunking logic - not a hard violation but suboptimal correctness. Response 1 is cleaner and more clearly correct. | google/gemini-3.1-pro-preview: tie (flipped on swap) — Both responses correctly implement the function and provide the required test cases. Response 2 is slightly better as it includes a helpful docstring, input validation, and a more concise, Pythonic implementation. / Both responses provide correct, well-structured Python functions that perfectly meet all requirements. They both correctly handle the edge case of an invoice exceeding the max_total and include the required 3 assert test cases. | x-ai/grok-4.5: tie — Both correctly implement ordered batching with oversized handling and supply 3 working asserts. Equally clear, complete, and useful; no task constraints broken. / Both correctly implement ordered batching with oversized handling and 3 assert tests. Equivalent quality, clarity, and usefulness for the task.
tieExplain codeCoding · 9 v 9Both correct and clear. Response 1 stays concise and under 150 words with one clear risk as asked. Response 2 exceeds scope slightly by adding…
Explain to a junior developer, in under 150 words, what this code does and one risk of using it: const cache = new Map(); function memo(fn) { return (...args) => { const k = JSON.stringify(args); if (!cache.has(k)) cache.set(k, fn(...args)); return cache.get(k); }; }
This code implements **memoization**, an optimization technique that caches the results of expensive function calls. When you wrap a function with `memo`, it returns a new function. This new function takes your arguments, converts them into a string key using `JSON.stringify`, and checks a global `Map`. If the result for those arguments is already cached, it returns it instantly. If not, it runs the original function, saves the result, and returns it. **One major risk:** Because the `cache` is a global variable that never clears, it can cause a **memory leak**. If you call the memoized function with many different arguments over time, the cache will grow indefinitely, eventually slowing down or crashing your application by exhausting available memory.
This code creates a simple memoization wrapper: it remembers function results so repeated calls with the same arguments do not run the original function again. - `cache` stores results in a `Map`. - `memo(fn)` returns a new function. - That new function converts its arguments to a JSON string key. - If no cached value exists, it calls `fn(...args)` and saves the result. - It then returns the saved result. Example: wrapping an expensive calculation can make repeated calls faster. One risk: the cache is global and never removes entries, so calling it with many different arguments can grow memory usage indefinitely. Also, `JSON.stringify` is not reliable for every value (for example, functions, `undefined`, circular objects, or objects with properties in different orders).
anthropic/claude-sonnet-5: tie (flipped on swap) — Both correct and clear. Response 1 stays concise and under 150 words with one clear risk as asked. Response 2 exceeds scope slightly by adding extra risk beyond 'one risk' instruction, though still under word count. / Both explain memoization correctly and note the memory leak risk. Response 1 also mentions JSON.stringify limitations, adding useful nuance, while staying within word limit. Response 2 is clear but slightly less thorough on risks. | google/gemini-3.1-pro-preview: tie (flipped on swap) — Both responses accurately explain the code and stay under the word limit. Response 1 is slightly better as it strictly adheres to providing exactly 'one risk' and explains the memory leak concept very clearly for a junior developer. / Both responses accurately explain the code and identify a valid risk within the word limit. Response 1 is slightly better structured with bullet points and mentions the JSON.stringify edge cases, which is highly valuable context for a junior developer. | x-ai/grok-4.5: tie — Both accurately explain memoization and the memory-leak risk in clear junior-friendly language under 150 words. Equally correct, compliant, useful and concise. / Both accurately explain memoization and the unbounded cache risk under 150 words, suitable for juniors. Equally clear, correct, and useful; minor extras in each balance out.
tieSQL queryCoding · 9.3 v 8.7Both queries are correct and nearly identical, using window functions properly. Response 1 includes tie-breaking by name in the RANK() ordering…
Given tables orders(id, customer_id, created_at, total) and customers(id, name, country), write a single SQL query returning each country's top 3 customers by lifetime spend in 2025, with columns country, name, total_spend, rank. Use a window function. Standard PostgreSQL.
```sql WITH customer_spend AS ( SELECT c.country, c.name, SUM(o.total) AS total_spend FROM orders o JOIN customers c ON c.id = o.customer_id WHERE o.created_at >= '2025-01-01'::timestamp AND o.created_at < '2026-01-01'::timestamp GROUP BY c.country, c.id, c.name ), ranked AS ( SELECT country, name, total_spend, RANK() OVER ( PARTITION BY country ORDER BY total_spend DESC, name ) AS rank FROM customer_spend ) SELECT country, name, total_spend, rank FROM ranked WHERE rank <= 3 ORDER BY country, rank, name; ``` Notes: - `RANK()` is used so ties share a rank; if you strictly want exactly three rows per country regardless of ties, replace it with `ROW_NUMBER()`. - The `created_at` filter uses timestamp casts to handle both `date` and `timestamp` column types safely. - `c.id` is included in the `GROUP BY` to disambiguate customers with identical names within the same country.
```sql WITH customer_spend AS ( SELECT c.country, c.name, c.id AS customer_id, SUM(o.total) AS total_spend FROM customers c JOIN orders o ON o.customer_id = c.id WHERE o.created_at >= DATE '2025-01-01' AND o.created_at < DATE '2026-01-01' GROUP BY c.country, c.name, c.id ), ranked_customers AS ( SELECT country, name, total_spend, RANK() OVER ( PARTITION BY country ORDER BY total_spend DESC ) AS rank FROM customer_spend ) SELECT country, name, total_spend, rank FROM ranked_customers WHERE rank <= 3 ORDER BY country, rank, name; ```
anthropic/claude-sonnet-5: tie (flipped on swap) — Both queries are correct and nearly identical, using window functions properly. Response 1 includes tie-breaking by name in the RANK() ordering, slightly more robust, and adds helpful notes explaining design choices. / Both queries correctly implement the task using window functions, proper joins, date filtering, and grouping. Both are clear, correct, and nearly identical in structure. Response 2 adds helpful notes but is otherwise functionally equivalent. | google/gemini-3.1-pro-preview: A — Both responses provide accurate and well-structured SQL queries that meet all requirements. Response 1 is slightly better because it includes a tie-breaker in the window function's ORDER BY clause, ensuring consistent ranking, and provides helpful explanatory notes. / Both responses provide accurate and well-structured SQL queries that meet all requirements. Response 2 is slightly better because it includes a deterministic sort in the window function (ordering by name as a tie-breaker) and provides helpful explanatory notes. | x-ai/grok-4.5: tie (flipped on swap) — Both correct SQL with window fn for top-3 2025 spend by country. R1 slightly better on ties/grouping/notes for clarity/usefulness. / Both are correct, complete PostgreSQL queries using RANK as required, with proper 2025 filtering and grouping. Equally clear, useful, and instruction-compliant; minor stylistic differences only.
tieRegexCoding · 8.2 v 8.8Both provide valid regex, JS one-liner, and edge case note as requested. Response 2 is more accurate/rigorous per official UK postcode spec, while…
Write a regex that matches UK postcodes like 'SW1A 1AA', 'M1 1AE', 'B33 8TH' (allow lowercase and optional space), and a one-line JavaScript example using it to validate a form field. Briefly note one edge case your regex does NOT handle.
**Regex:** ```javascript /^[a-z]{1,2}\d[a-z\d]? ?\d[a-z]{2}$/i ``` *(The `i` flag makes it case-insensitive, allowing both uppercase and lowercase letters).* **JavaScript one-liner:** ```javascript const isValid = /^[a-z]{1,2}\d[a-z\d]? ?\d[a-z]{2}$/i.test(document.querySelector('#postcode').value); ``` **Edge case not handled:** This regex does not handle the special historic postcode **"GIR 0AA"** (used by the National Girobank), nor does it enforce strict geographical rules, such as the fact that the first letter of a UK postcode is never 'Q', 'V', or 'X'.
```regex /^(?:GIR ?0AA|(?:[A-PR-UWYZ][0-9][0-9]?|[A-PR-UWYZ][A-HK-Y][0-9][0-9]?|[A-PR-UWYZ][0-9][A-HJKS-UW]) ?[0-9][ABD-HJLNP-UW-Z]{2})$/i ``` ```js const valid = /^(?:GIR ?0AA|(?:[A-PR-UWYZ][0-9][0-9]?|[A-PR-UWYZ][A-HK-Y][0-9][0-9]?|[A-PR-UWYZ][0-9][A-HJKS-UW]) ?[0-9][ABD-HJLNP-UW-Z]{2})$/i.test(document.querySelector('#postcode').value); ``` This accepts lowercase and either no space or one space before the final three characters. One edge case not handled is **BFPO** (British Forces Post Office) addresses, such as `BFPO 1`.
anthropic/claude-sonnet-5: tie (flipped on swap) — Both provide valid regex, JS one-liner, and edge case note as requested. Response 2 is more accurate/rigorous per official UK postcode spec, while Response 1 is simpler but less precise. Both fulfill instructions equally well overall. / Both meet requirements; Response 1's regex is more precise (handles GIR 0AA and stricter letter rules), while Response 2's is simpler but less accurate and admits it misses GIR 0AA. Both clear and concise. | google/gemini-3.1-pro-preview: tie (flipped on swap) — Response 1 provides a simpler, more readable regex that perfectly balances the requirements while remaining accessible to a non-technical professional. Response 2 provides an overly complex regex (the official strict version) which is harder to read and maintain. / Response 1 provides a highly accurate, standard regex for UK postcodes, whereas Response 2 provides a naive regex that allows many invalid letter combinations. Both follow all instructions perfectly. | x-ai/grok-4.5: B — Both meet all requirements with working regex+JS+edge note. R2 more factually complete/accurate for real UK formats; R1 simpler but looser. Equal clarity/usefulness. / Resp1 has highly accurate UK postcode regex covering formats/GIR; Resp2 is simpler/looser but works for examples. Both follow all task parts clearly and usefully.
tieMulti-step logicReasoning · 8.3 v 7.2Response 1 correctly treats the £60 as total extra labour cost, yielding £209/week profit. Response 2 incorrectly scales labour per loaf…
A bakery sells loaves at £3.20. Ingredients cost £1.10/loaf, labour £0.90/loaf, fixed costs £480/week. They sell 450 loaves/week. A supplier offers ingredients at £0.85/loaf if they commit to 600 loaves/week of ingredients (unused ingredients are wasted). Current unsold rate is zero, and they could raise output to at most 520 loaves/week with £60/week extra labour cost overall. Should they take the deal? Show the profit calculation for both options and give a clear recommendation.
Here is the step-by-step profit calculation for both options to help you make a decision. ### Option 1: Current Situation Currently, you are producing and selling 450 loaves a week with an unsold rate of zero. **Revenue:** * 450 loaves × £3.20 = **£1,440** **Costs:** * Ingredients: 450 loaves × £1.10 = £495 * Labour: 450 loaves × £0.90 = £405 * Fixed costs: £480 * **Total Costs = £495 + £405 + £480 = £1,380** **Current Profit:** £1,440 (Revenue) - £1,380 (Costs) = **£60 / week** --- ### Option 2: Take the Supplier Deal (Increase output to 520 loaves) If you take the deal, you must buy ingredients for 600 loaves, but you can only produce and sell a maximum of 520 loaves (meaning 80 loaves worth of ingredients will be wasted). To produce these extra 70 loaves, your overall labour costs will increase by £60. **Revenue:** * 520 loaves × £3.20 = **£1,664** **Costs:** * Ingredients: 600 loaves (committed) × £0.85 = £510 * Labour: £405 (current) + £60 (extra overall) = £465 * Fixed costs: £480 * **Total Costs = £510 + £465 + £480 = £1,455** **New Profit:** £1,664 (Revenue) - £1,455 (Costs) = **£209 / week** *(Note: If you took the deal but kept output at 450 loaves, you would waste 150 loaves of ingredients. Your costs would be £1,395 against £1,440 revenue, yielding only £45/week. Therefore, if you take the deal, you must increase output to 520 loaves).* --- ### Recommendation **Yes, you should take the deal.** By committing to the supplier's offer and raising your output to the maximum capacity of 520 loaves per week, your weekly profit will increase from **£60 to £209**. Even though you will be wasting ingredients for 80 loaves, the significantly cheaper ingredient price and the additional sales from the extra 70 loaves easily outweigh the £60 extra labour cost and the wasted stock.
### Option 1: Keep current arrangement (450 loaves/week) - Revenue: \(450 \times £3.20 = £1,440\) - Ingredients: \(450 \times £1.10 = £495\) - Labour: \(450 \times £0.90 = £405\) - Fixed costs: £480 \[ \text{Profit} = £1,440 - £495 - £405 - £480 = \boxed{£60\text{/week}} \] --- ### Option 2: Take supplier deal and increase production to 520 loaves/week The bakery must buy ingredients for 600 loaves: - Ingredient cost: \(600 \times £0.85 = £510\) - Ingredients wasted: \(600 - 520 = 80\) loaves’ worth - Revenue: \(520 \times £3.20 = £1,664\) - Labour: \(520 \times £0.90 = £468\), plus £60 extra labour cost \[ £468 + £60 = £528 \] - Fixed costs: £480 \[ \text{Profit} = £1,664 - £510 - £528 - £480 = \boxed{£146\text{/week}} \] --- ### Comparison | Option | Weekly profit | |---|---:| | Current arrangement, 450 loaves | £60 | | Supplier deal, 520 loaves | £146 | \[ £146 - £60 = \boxed{£86\text{ more profit per week}} \] ### Recommendation **Yes — take the supplier deal, provided the bakery can sell all 520 loaves each week.** Even after wasting ingredients for 80 loaves and paying the additional labour cost, weekly profit rises from **£60 to £146**, an increase of **£86 per week**.
anthropic/claude-sonnet-5: tie (flipped on swap) — Response 1 correctly treats the £60 as total extra labour cost, yielding £209/week profit. Response 2 incorrectly scales labour per loaf (520×0.90) plus adds £60, double-counting labour cost and understating profit (£146). Both are clear and well-structured, but Response 1 is more accurate per the stated constraint. / Response 1's calculations are correct (£146 profit). Response 2 incorrectly uses labour £405 instead of 520×£0.90=£468, understating costs and inflating profit to £209, a factual error despite correct recommendation direction. | google/gemini-3.1-pro-preview: A — Response 1 correctly interprets the '£60 extra labour cost overall' as the total additional labour cost for the extra production. Response 2 incorrectly calculates labour by applying the per-loaf rate to the new total and then adding £60 on top. / Response 2 correctly interprets '£60/week extra labour cost overall' as a flat £60 addition to the current total labour cost. Response 1 incorrectly applies the per-loaf labour rate to the new volume and then adds the £60 on top, double-counting the labour for the extra loaves. | x-ai/grok-4.5: tie (flipped on swap) — R1 correctly computes labour as 405+60=465 and profit 209. R2 wrongly uses 520*0.90+60=528, understating profit at 146. Both clear and recommend taking the deal; R1 more accurate/useful. / R1 correctly applies £0.90/loaf labour +£60 extra (profit £146). R2 undercounts labour as only +£60 on base 450 (wrong £209). Both recommend correctly and structure well; R1 clearer.
tiePlanningReasoning · 8.5 v 8.3Both meet format: one line per day with owner, riskiest dependency flagged. Response 2 slightly more detailed and specific with budget allocation…
Plan the launch week for a small online course (already recorded). Resources: one founder, a part-time VA (10h), email list of 2,000, £300 ad budget. Produce a 7-day plan, one line per day, each line naming the owner. Flag the single riskiest dependency.
**7-Day Launch Plan** Day 1: Founder sends the "coming soon" teaser email to the 2k list while VA sets up and launches the £300 ad campaign. Day 2: Founder publishes a value-packed mini-lesson on socials while VA monitors ad spend and replies to comments. Day 3: Founder sends the "doors are open" email with the checkout link while VA tests the checkout flow and tags non-openers. Day 4: Founder hosts a 30-minute live Q&A to drive sales while VA repurposes the session into short clips for ad retargeting. Day 5: Founder sends an email tackling FAQs and objections while VA updates ad copy to emphasize the limited-time offer. Day 6: Founder sends the "24 hours left" urgency email while VA sends manual DMs to highly engaged leads who haven't bought. Day 7: Founder sends the final "last chance" and "closing in 2 hours" emails while VA pauses ads and closes the cart at midnight. **Riskiest Dependency:** The VA's strict 10-hour weekly limit, which could easily be exhausted by Day 4 if ad setup or comment moderation takes longer than expected, leaving Days 5-7 without support.
- **Day 1 — Founder:** Open enrolment, send launch email #1 to the 2,000-person list, publish sales page, and announce on personal/social channels. - **Day 2 — VA:** Check checkout, links, tracking, and inbox; compile FAQs/objections from replies; Founder records short FAQ post/video. - **Day 3 — Founder:** Send value-led email #2 (lesson/story + course outcome), host a short live Q&A, and invite attendees to enrol. - **Day 4 — VA:** Launch £100 retargeting-focused ad campaign (site visitors/email-engagers where possible), monitor spend and report leads/sales to Founder. - **Day 5 — Founder:** Send social-proof/case-study email #3, publish testimonials or behind-the-scenes content, and personally reply to high-intent prospects. - **Day 6 — VA:** Increase or redirect remaining £200 ad spend to best-performing creative/audience; schedule final-day emails and social reminders. - **Day 7 — Founder:** Send “closes today” morning email and final-hours email, post deadline reminders, run a final live/recorded Q&A, then close enrolment and thank buyers. **Riskiest dependency:** The sales page and checkout must be fully working before Day 1; a broken or unclear purchase flow would waste the email-list attention and limited ad budget.
anthropic/claude-sonnet-5: tie — Both meet format: one line per day with owner, riskiest dependency flagged. Response 2 slightly more detailed and specific with budget allocation; Response 1 simpler but clear. Both accurate and useful, comparable clarity. / Both provide clear 7-day plans with owners and a risk flag. Response 2 checks checkout flow on Day3 rather than before launch, slightly less safe; Response 1 flags checkout risk more sensibly. Both similar quality, clarity, usefulness. | google/gemini-3.1-pro-preview: A — Response 1 provides a highly practical plan that correctly incorporates all constraints. Its identification of the VA's 10-hour limit as the riskiest dependency is exceptionally insightful for this specific scenario, whereas Response 2 offers a more generic dependency. / Both responses follow all instructions. Response 2 is slightly better because its flagged dependency (the VA's 10-hour limit) is highly specific to the provided constraints and very insightful for a busy professional. | x-ai/grok-4.5: tie (flipped on swap) — Both deliver valid 7-day one-line plans naming owners plus one risk. R1 is tighter, more concise and better matches the exact format while staying realistic on the 10h VA limit. / Both meet format (7 one-line days naming owners + risk flag). R1 more realistic sequencing and ad use; R2 front-loads full budget and teaser poorly. R1 clearer/more useful.
tieTrade-off analysisReasoning · 8.8 v 9Both stay under 200 words and cover all elements. Response 2 gives more precise, nuanced cost figures and threshold analysis, plus a sharper…
A 12-person agency must choose: (A) hire a mid-level developer at £55k, or (B) contract overflow work to freelancers at roughly £400/day, expected 60 days/year. Give a recommendation in under 200 words covering cost, flexibility, quality risk, and one non-obvious factor.
I recommend **Option B (Freelancers)**. **Cost:** A £55k salary balloons to roughly £70k with taxes and benefits, plus fixed overhead. Freelancers cost exactly £24k for 60 days (£400/day), saving the agency approximately £46k annually. **Flexibility:** Option B wins decisively. If the overflow work dries up, you simply stop contracting. Option A locks you into a fixed monthly burn rate, risking painful layoffs if the pipeline slows. **Quality Risk:** Option B carries higher risk due to onboarding friction and a lack of institutional knowledge. Mitigate this by building a vetted roster of 2-3 trusted freelancers rather than hiring ad-hoc per project. **Non-obvious factor:** Opportunity cost. Reinvesting the £46k annual savings into targeted sales and marketing can generate enough sustained pipeline to turn temporary overflow into permanent growth. This approach eventually justifies the £55k hire with a guaranteed full workload, rather than gambling on the hire first and hoping the work continues.
**Recommendation: Choose B — freelancers — for the current demand level.** At 60 days/year, freelancers cost about **£24k/year** (plus VAT where applicable). A £55k employee will likely cost **£65–75k fully loaded** once employer NI, pension, equipment, software, recruitment and management time are included. Freelancers provide strong **flexibility**: scale capacity up or down, access specialist skills, and avoid carrying fixed cost during quieter periods. Hiring becomes financially compelling only if demand is sustained—roughly **160–180 contractor days/year** depending on the employee’s fully loaded cost. The main downside is **quality and continuity risk**: external developers need onboarding, may lack product context, and can create inconsistent standards or dependency on individuals. Mitigate with clear specs, code review, documentation, repository access, and a preferred supplier pool. A non-obvious factor: **management bandwidth**. In a 12-person agency, repeatedly briefing and supervising freelancers can consume senior staff time; if that starts delaying delivery or sales, a permanent hire may be worth the premium earlier than the spreadsheet suggests.
anthropic/claude-sonnet-5: tie (flipped on swap) — Both stay under 200 words and cover all elements. Response 2 gives more precise, nuanced cost figures and threshold analysis, plus a sharper non-obvious factor (management bandwidth). Response 1 is solid but less precise/quantitative. / Both correctly recommend B, cover cost/flexibility/quality/non-obvious factor concisely under 200 words with clear structure. Similar quality; Response 2 slightly more structured with headers, Response 1 slightly more nuanced on threshold analysis. Essentially equivalent. | google/gemini-3.1-pro-preview: tie — Both responses perfectly follow all instructions, staying under the 200-word limit while covering all required points. They offer excellent, practical advice for a small agency, with Response 1 highlighting opportunity cost and Response 2 highlighting management bandwidth. Both are highly useful and concise. / Both responses perfectly follow all instructions, including the word count constraint. They accurately address cost, flexibility, quality risk, and provide excellent non-obvious factors tailored to a small agency. Both are highly useful and concise. | x-ai/grok-4.5: tie — Both accurately compare costs (~£24k vs ~£70k), flexibility, quality risks with mitigations, and a solid non-obvious factor. Both under 200 words, clear, structured, and useful for non-technical readers. Equal overall. / Both under 200 words, cover all required topics accurately with solid costs, risks, and non-obvious factors. Equally clear, useful, and instruction-compliant for a busy pro.
tieMaths word problemReasoning · 9.7 v 9.7Both correctly compute 10:25 with identical clear working. Response 1 uses more accessible formatting for a non-technical audience, while Response…
A tank holds 2,400 litres. Pump A fills at 40 L/min. Pump B drains at 25 L/min. A runs from 09:00. B accidentally switches on at 09:20. At 10:00 B is switched off. At what time is the tank full? Show your working.
Here is the step-by-step working to find out when the tank will be full: **1. From 09:00 to 09:20 (Only Pump A is running)** * **Duration:** 20 minutes * **Fill rate:** 40 L/min (Pump A) * **Water added:** 20 minutes × 40 L/min = **800 litres** * **Total in tank at 09:20:** 800 litres **2. From 09:20 to 10:00 (Both Pump A and Pump B are running)** * **Duration:** 40 minutes * **Net fill rate:** 40 L/min (Pump A) - 25 L/min (Pump B) = 15 L/min * **Water added:** 40 minutes × 15 L/min = **600 litres** * **Total in tank at 10:00:** 800 litres + 600 litres = **1,400 litres** **3. From 10:00 onwards (Only Pump A is running)** * **Remaining capacity:** 2,400 litres (total) - 1,400 litres (current) = **1,000 litres needed** * **Fill rate:** 40 L/min (Pump A) * **Time needed to fill the rest:** 1,000 litres ÷ 40 L/min = **25 minutes** **Conclusion:** Add the 25 minutes needed to the time Pump B was switched off (10:00). 10:00 + 25 minutes = **10:25** The tank will be full at **10:25**.
Assuming the tank is empty at 09:00: - From 09:00 to 09:20, only Pump A runs: \(20 \times 40 = 800\) L filled. - From 09:20 to 10:00, both pumps run. Net filling rate: \(40 - 25 = 15\) L/min. Over 40 minutes: \(40 \times 15 = 600\) L filled. - Amount in tank at 10:00: \(800 + 600 = 1400\) L. - Remaining capacity: \(2400 - 1400 = 1000\) L. - After 10:00, Pump B is off, so Pump A fills at 40 L/min: \[ 1000 \div 40 = 25 \text{ minutes} \] Therefore, the tank is full at **10:25**.
anthropic/claude-sonnet-5: tie — Both correctly compute 10:25 with identical clear working. Response 1 uses more accessible formatting for a non-technical audience, while Response 2 uses LaTeX which is slightly less friendly. Both are correct and concise. / Both responses correctly compute the same answer (10:25) with identical correct working. Response 2 is slightly more clearly formatted but longer; Response 1 is more concise. Both fully correct and equally clear overall. | google/gemini-3.1-pro-preview: tie (flipped on swap) — Both responses are perfectly accurate and show clear working. Response 1 is slightly better formatted for a non-technical professional, avoiding mathematical notation in favor of plain text explanations. / Both responses correctly solve the math problem with clear, step-by-step working. They are both accurate, easy to follow, and perfectly address the prompt. | x-ai/grok-4.5: tie — Both correctly compute 800L + 600L = 1400L by 10:00, then 25 min more to fill remaining 1000L at 40 L/min, reaching 10:25. Clear working, fully useful, no constraints broken. / Both correctly compute 800+600=1400L by 10:00 then +25min at 40L/min to reach 2400L at 10:25. Clear working, no constraints broken.
tieCatch the flawReasoning · 9.2 v 9.2Both correctly identify causal and linearity flaws, under 100 words, clear and useful. Response 2 adds actionable suggestion (controlled test)…
Find the flaw in this argument and explain it in under 100 words: 'Our churn dropped from 6% to 3% after we introduced the loyalty programme in March, so the programme cut churn in half. We should double the loyalty budget to cut churn to 1.5%.'
The argument commits two flaws: **post hoc fallacy** and **false linearity**. First, it assumes the loyalty programme caused the churn drop, ignoring other factors like seasonality or market shifts. Second, it assumes halving churn again requires just doubling the budget. Diminishing returns make this highly unlikely; the easiest customers to retain are already kept, so cutting churn further will require disproportionately more effort and money, not a simple linear budget increase.
The argument confuses correlation with causation: churn may have fallen because of seasonality, pricing changes, customer mix, or other factors—not necessarily the loyalty programme. It also assumes a linear, proportional effect: even if the programme caused some reduction, doubling spend may produce diminishing returns or no further reduction. A controlled test or comparison group is needed before attributing the drop or increasing budget.
anthropic/claude-sonnet-5: tie — Both correctly identify causal and linearity flaws, under 100 words, clear and useful. Response 2 adds actionable suggestion (controlled test), slightly more useful; Response 1 slightly more structured. Essentially equivalent quality. / Both correctly identify correlation/causation and linearity flaws concisely, under 100 words, no constraint violations. Response 2 is slightly more structured and detailed; Response 1 slightly more concise. Both equally useful and clear. | google/gemini-3.1-pro-preview: tie (flipped on swap) — Both responses correctly identify the flaws (causation vs. correlation and linear extrapolation) and stay well under the 100-word limit. Response 1 is slightly better structured with bold terms and bullet-like paragraphs, making it easier for a busy professional to digest quickly. / Both responses correctly identify the two flaws (correlation vs causation and linear scaling/diminishing returns) and stay well under the 100-word limit. Response 1 is slightly better for a non-technical professional as it avoids academic jargon and includes a practical next step (controlled test). | x-ai/grok-4.5: tie — Both correctly identify post-hoc causation and false linearity/diminishing returns flaws, stay under 100 words, and are clear, concise, and useful for non-technical readers. / Both correctly identify correlation-causation and non-linearity flaws, stay under 100 words, and give clear useful explanations for non-experts.
tieMessy to JSONExtraction · 9.3 v 9.3Both responses correctly extract all fields into valid JSON with accurate information. Formatting differs slightly (indented vs compact) but both…
Extract to JSON with keys name, company, phone, email, intent: 'hiya - jules moreau here from Atelier Nord (the lighting people). best number is 07700 900 457, or jules@ateliernord.fr. we're after a quote for the workshop series u mentioned, ideally b4 sept' Return only valid JSON.
{ "name": "Jules Moreau", "company": "Atelier Nord", "phone": "07700 900 457", "email": "jules@ateliernord.fr", "intent": "Quote for the workshop series before September" }
{"name":"Jules Moreau","company":"Atelier Nord","phone":"07700 900 457","email":"jules@ateliernord.fr","intent":"Requesting a quote for the workshop series mentioned, ideally before September."}
anthropic/claude-sonnet-5: tie — Both responses correctly extract all fields into valid JSON with accurate information. Formatting differs slightly (indented vs compact) but both meet the 'return only valid JSON' constraint equally well. Intent phrasing differs marginally but both are accurate and clear. / Both correctly extract fields into valid JSON with accurate values. Response 1 has slightly more complete intent phrasing, but both are functionally equivalent and equally clear and concise. | google/gemini-3.1-pro-preview: tie — Both responses perfectly followed the instructions, accurately extracted the required information, and returned only valid JSON. Response 1 is pretty-printed while Response 2 is minified, but both are equally correct and useful. / Both responses perfectly follow the instructions, extracting the correct information into valid JSON with the specified keys. They are both accurate and useful. | x-ai/grok-4.5: tie — Both extract all fields accurately into valid JSON only. Minor intent phrasing differences; equally correct, compliant, useful, and clear. / Both extract accurate fields into valid JSON only, matching all keys and source data with clear intents. Equally correct, compliant, useful and concise.
GPT-5.6 TerraTable from proseExtraction · 9 v 9.5Both correctly compute percentages and follow instructions; Response 2 adds right-alignment and a plus sign, minor stylistic improvements but not…
Turn this into a markdown table with columns Product, Q1, Q2, Change: 'The Alpha line did 14,200 units in Q1 and 15,800 in Q2. Beta slipped from 9,100 to 7,400. The new Gamma launched mid-Q2 with 3,300 units (no Q1 sales). Delta held flat at 5,000 both quarters.' Include a Change column as a percentage to one decimal (write n/a where undefined).
| Product | Q1 | Q2 | Change | |---|---|---|---| | Alpha | 14,200 | 15,800 | 11.3% | | Beta | 9,100 | 7,400 | -18.7% | | Gamma | 0 | 3,300 | n/a | | Delta | 5,000 | 5,000 | 0.0% |
| Product | Q1 | Q2 | Change | |---|---:|---:|---:| | Alpha | 14,200 | 15,800 | +11.3% | | Beta | 9,100 | 7,400 | -18.7% | | Gamma | 0 | 3,300 | n/a | | Delta | 5,000 | 5,000 | 0.0% |
anthropic/claude-sonnet-5: tie (flipped on swap) — Both correctly compute percentages and follow instructions; Response 2 adds right-alignment and a plus sign, minor stylistic improvements but not required. Both equally accurate and clear. / Both nearly identical and accurate; Response 1 uses consistent +/- sign formatting for clarity, slightly better than Response 2 which omits + for Alpha. | google/gemini-3.1-pro-preview: B — Both responses follow all instructions perfectly and calculate the percentages correctly. Response 2 is slightly better because it right-aligns the numerical columns, which is standard best practice for tables containing numbers. / Both responses follow all instructions perfectly and calculate the percentages correctly. Response 1 is slightly better because it right-aligns the numerical columns and includes a plus sign for the positive change, which improves readability. | x-ai/grok-4.5: B — Both accurate with correct % calcs and n/a. Resp2 superior alignment and + sign aid clarity/usefulness for pros; both follow format exactly. / Both correct tables with accurate % (Alpha +11.3, Beta -18.7, Gamma n/a, Delta 0.0). R1 better with + sign and right-align; R2 omits + on positive change.
GLM 5.2Contact detailsExtraction · 8.8 v 7.7Both format correctly, but Response 1 correctly assigns Priya Shah's role as 'Legal sign-off' while Response 2 marks it unknown, losing…
List every person mentioned below with their role and organisation, one line each in the format Name - Role - Org. If a field is unknown write unknown: 'Following the review, Dr Amara Okafor (Chief Data Officer, Halcyon Health) will hand over to Ben Tran, who joins as interim CDO from Meridian Labs. Legal sign-off sits with Priya Shah at Halcyon; the external audit remains with T. Lindqvist.'
Dr Amara Okafor - Chief Data Officer - Halcyon Health Ben Tran - interim CDO - Meridian Labs Priya Shah - Legal sign-off - Halcyon T. Lindqvist - External audit - unknown
Dr Amara Okafor - Chief Data Officer - Halcyon Health Ben Tran - interim CDO - Meridian Labs Priya Shah - unknown - Halcyon T. Lindqvist - external audit - unknown
anthropic/claude-sonnet-5: tie (flipped on swap) — Both format correctly, but Response 1 correctly assigns Priya Shah's role as 'Legal sign-off' while Response 2 marks it unknown, losing information. Minor ambiguity on Ben Tran's org unaddressed by both. / Both responses correctly list all four individuals with roles/orgs in the required format; nearly identical content, minor differences in role labeling (unknown vs inferred) but both reasonable. | google/gemini-3.1-pro-preview: A — Response 1 correctly identifies Priya Shah's role based on the text ('Legal sign-off'), whereas Response 2 incorrectly lists it as 'unknown'. Both follow the formatting constraints perfectly. / Both responses follow the formatting constraints perfectly. Response 2 is slightly better as it extracts the implied roles/responsibilities (Legal sign-off, External audit) rather than defaulting to 'unknown' for Priya Shah. | x-ai/grok-4.5: A — Both extract the four people accurately in required format. R1 better assigns Priya's implied legal role; R2 leaves it unknown. Minor org/role ambiguities for Ben and Lindqvist shared. / R2 correctly extracts Priya Shah role as Legal sign-off; R1 wrongly marks it unknown. Both match required format and list all four people accurately otherwise.
GLM 5.2Key factsExtraction · 9.3 v 6.8Response 1 covers all 5 distinct facts within word limits. Response 2 splits phase one into two bullets, omitting the construction traffic fact…
Extract exactly 5 key facts as bullets (each under 15 words) from: 'The council approved the harbour redevelopment on 12 June by 7 votes to 2. Phase one, costing £4.6m, begins in October and includes the fish market roof. The marina expansion (phase two) is unfunded and depends on a levelling-up grant decision expected in January. Local traders get first refusal on the 14 new retail units. Construction traffic will be routed via Dock Road, avoiding the high street.'
* Council approved the harbour redevelopment on 12 June by 7-2. * Phase one costs £4.6m, begins in October, and includes the fish market roof. * The unfunded phase two marina expansion depends on a January grant decision. * Local traders get first refusal on the 14 new retail units. * Construction traffic will be routed via Dock Road to avoid the high street.
- Council approved harbour redevelopment on 12 June by 7 votes to 2. - Phase one costs £4.6m and begins in October. - Phase one includes replacing the fish market roof. - Marina expansion depends on a levelling-up grant decision expected in January. - Local traders receive first refusal on 14 new retail units.
anthropic/claude-sonnet-5: A — Response 1 covers all 5 distinct facts within word limits. Response 2 splits phase one into two bullets, omitting the construction traffic fact, failing to capture all key facts. / Both meet the 5-bullet, <15-word constraint. Response 2 captures all 5 distinct facts including the traffic routing detail, which Response 1 omits, making it more complete and useful. | google/gemini-3.1-pro-preview: A — Both responses follow all constraints, including exactly 5 bullets under 15 words each. Response 1 is better because it captures a wider range of key facts from the source text, whereas Response 2 splits one fact into two and misses the construction traffic detail. / Both responses follow all constraints perfectly. Response 2 is better because it captures a broader range of the original text's information, including the construction traffic detail, whereas Response 1 splits one sentence into two bullets and misses the final fact. | x-ai/grok-4.5: A — Resp1 accurately extracts all 5 key facts within limits. Resp2 adds unstated 'replacing', omits traffic routing and unfunded status, less complete. / R2 accurately extracts all 5 key facts within limits; R1 invents 'replacing', omits traffic routing, less complete.
tieSchema complianceExtraction · 9.7 v 9.7Both correctly match the schema and data, differing only in formatting whitespace. Both fully satisfy accuracy and instructions.
Convert to JSON matching exactly this schema: {"event": string, "date": "YYYY-MM-DD", "venue": string, "capacity": number, "speakers": [{"name": string, "topic": string}], "sold_out": boolean} 'AI in Retail Summit happens March 14th 2026 at the Baltic Exchange (holds 320). Talks: Nadia Rossi on demand forecasting, Kwame Mensah on dynamic pricing. Tickets still available.' Return only the JSON.
{ "event": "AI in Retail Summit", "date": "2026-03-14", "venue": "Baltic Exchange", "capacity": 320, "speakers": [ { "name": "Nadia Rossi", "topic": "demand forecasting" }, { "name": "Kwame Mensah", "topic": "dynamic pricing" } ], "sold_out": false }
{"event":"AI in Retail Summit","date":"2026-03-14","venue":"Baltic Exchange","capacity":320,"speakers":[{"name":"Nadia Rossi","topic":"demand forecasting"},{"name":"Kwame Mensah","topic":"dynamic pricing"}],"sold_out":false}
anthropic/claude-sonnet-5: tie — Both correctly match the schema and data, differing only in formatting whitespace. Both fully satisfy accuracy and instructions. / Both correctly match the schema with accurate data; only formatting differs (compact vs pretty-printed), which is immaterial to correctness. | google/gemini-3.1-pro-preview: tie — Both responses perfectly follow the instructions, accurately extracting the information into the exact JSON schema requested, and returning only the JSON. Response 1 is pretty-printed while Response 2 is minified, but both are completely valid. / Both responses perfectly extract the information into the requested JSON schema and follow the negative constraint to return only the JSON. Response 1 is minified while Response 2 is pretty-printed, but both are equally valid. | x-ai/grok-4.5: tie — Both produce exact valid schema-matching JSON with correct data extraction and only JSON output. No differences in accuracy or compliance. / Both produce exact valid schema-matching JSON with correct data extraction and only JSON output. No constraints violated.
tieFormat constraintsInstruction-following · 9.2 v 8.8Both meet sentence count, starting words, and word limits. R1 is slightly clearer and more accurate in describing photosynthesis basics, while R2…
Describe how photosynthesis works in exactly 3 sentences. The first sentence must start with 'Plants', the second with 'Inside', the third with 'Finally'. No sentence may exceed 20 words.
Plants capture sunlight to transform water and carbon dioxide into energy. Inside the chloroplasts, chlorophyll absorbs light energy to drive chemical reactions. Finally, these reactions produce glucose for growth and release oxygen as a byproduct.
Plants capture sunlight with chlorophyll, using its energy to split water and energize electrons. Inside chloroplasts, electrons help convert carbon dioxide into energy-rich sugars while releasing oxygen. Finally, plants use sugars for growth, storage, and cellular respiration.
anthropic/claude-sonnet-5: tie (flipped on swap) — Both meet sentence count, starting words, and word limits. R1 is slightly clearer and more accurate in describing photosynthesis basics, while R2 is also correct but slightly denser in wording. / Both responses meet the 3-sentence structure, correct starting words, and stay within 20-word limits per sentence. Content is accurate and clear in both. Neither has a clear advantage in usefulness or clarity. | google/gemini-3.1-pro-preview: A — Both responses follow all constraints perfectly. Response 1 is slightly better for a non-technical professional as it avoids jargon like 'energize electrons', making it clearer and more accessible. / Both responses follow all constraints perfectly. Response 2 is slightly better suited for a non-technical professional as it avoids mentioning 'energize electrons', keeping the explanation simpler and more accessible while remaining accurate. | x-ai/grok-4.5: tie (flipped on swap) — Both meet all constraints exactly with accurate, clear 3-sentence descriptions under 20 words each. Equally useful and concise for non-experts. / Both meet all constraints exactly with accurate starts and word counts. R1 slightly more precise on mechanisms; both clear, useful, and concise for non-experts.
tieWord limitInstruction-following · 9.3 v 9Both responses are exactly 50 words, accurate, and clear. R2 adds a useful note on courts assessing substance, slightly more nuanced, but both are…
Explain the difference between a lease and a licence for UK commercial property in exactly 50 words. Count carefully - exactly 50.
A lease grants exclusive possession of UK commercial property for a fixed term, creating a legal proprietary interest that can be assigned and protected by statute. A licence merely provides personal permission to occupy without exclusive possession, creating no legal estate, offering less security, and preventing assignment to third parties.
A lease grants exclusive possession of premises for a term, usually creating a proprietary interest enforceable against third parties. A licence merely permits occupation or use, without exclusive possession or an estate in land, and is generally personal, revocable, and less protected. Courts assess substance, not labels, when distinguishing them.
anthropic/claude-sonnet-5: tie — Both responses are exactly 50 words, accurate, and clear. R2 adds a useful note on courts assessing substance, slightly more nuanced, but both are equally correct and concise. / Both responses are exactly 50 words, accurate, clear, and cover key distinctions (exclusive possession, proprietary interest vs personal permission). Content quality and clarity are comparable, making this a tie. | google/gemini-3.1-pro-preview: A — Both responses successfully meet the strict 50-word constraint and accurately explain the difference. Response 1 is slightly better as it explicitly references 'UK commercial property' as requested in the prompt. / Both responses successfully meet the strict 50-word constraint and accurately explain the legal distinction. Response 2 is slightly better as it explicitly references 'UK commercial property' and highlights statutory protection and assignability, which are highly relevant for commercial contexts. | x-ai/grok-4.5: tie — Both exactly 50 words, factually accurate on UK lease vs licence distinctions, clear and useful. Minor edge to 1 for task-specific phrasing but overall equal. / Both exactly 50 words, factually accurate on exclusive possession and proprietary vs personal rights, clear and useful. Minor edge to 1 for substance-over-form point, but overall equal.
GPT-5.6 TerraForbidden wordsInstruction-following · 8 v 8.7Response 1 uses 'caffeine'-like term 'stimulant' but also says 'over extraction' fine; actually check banned words: 'roast' variant 'cooked' fine…
Explain what makes coffee taste bitter, in 80-120 words, WITHOUT using any of these words: bitter, bean, roast, caffeine, acid. Do not use hyphenated or partial variants of them either.
Coffee's harsh, sharp flavor comes from a few chemical reactions during its preparation. When the raw seeds of the coffee fruit are cooked at high temperatures, complex compounds break down. The longer and darker they are cooked, the more these compounds transform into harsh tasting molecules. Additionally, the natural stimulant found in the seeds contributes a sharp edge to the flavor profile. Over extraction during brewing also pulls out heavy, astringent substances from the grounds. If the water is too hot or the brewing time is too long, these deep, heavy flavors overpower the sweeter, fruity notes, leaving a strong, sharp taste on the tongue.
Coffee’s sharp, harsh edge comes from several compounds released during heating and brewing. Longer, hotter preparation creates more dark, smoky molecules and breaks down natural sugars, reducing sweetness that might otherwise soften the cup. Fine grinding, very hot water, and an overly long contact time pull out heavier substances from the grounds, giving a dry, puckering finish. Water rich in certain minerals can emphasize this effect, while stale material may taste woody or ashy. Choosing a lighter preparation, using cooler water, shortening extraction, and adding milk or sugar can make the drink seem smoother.
anthropic/claude-sonnet-5: B — Response 1 uses 'caffeine'-like term 'stimulant' but also says 'over extraction' fine; actually check banned words: 'roast' variant 'cooked' fine, but 'stimulant' okay. Wait 'raw seeds' avoids 'bean'. Response 1 seems compliant too, but word count ~105 fine. Response 2 more concise, accurate, avoids banned words, slightly better clarity. / Both avoid banned words and fit word count. Response 1 is clearer, more concise, and offers practical tips; Response 2 is slightly wordier but still accurate and instructive. | google/gemini-3.1-pro-preview: A — Both responses follow all negative constraints and word count limits perfectly. Response 1 is slightly better focused on explaining the causes of the harsh taste, whereas Response 2 spends a significant portion on how to mitigate it. / Both responses follow all constraints, including word count and forbidden words. Response 2 provides a slightly more comprehensive explanation by cleverly alluding to the cooking process (roasting) and the natural stimulant (caffeine), which are primary contributors to the harsh flavor. | x-ai/grok-4.5: B — Both meet 80-120 words and avoid all banned terms/variants. Accurate explanations of heat, extraction, compounds. R2 slightly clearer, more useful tips without fluff. / Both meet word count and ban all forbidden terms/variants with accurate explanations. Resp1 edges out via practical tips and smoother flow for non-experts.
tieMulti-constraintInstruction-following · 9 v 9.3Both meet word count, 3 paragraphs, single question, phrase included, one emoji, no hashtags. Response 2 is slightly more engaging with a clear…
Write a LinkedIn post about hiring your first apprentice. Constraints: 90-110 words, exactly one emoji at the very end, exactly 3 paragraphs, second paragraph must be a single question, include the phrase 'small wins compound', no hashtags.
Today marks a huge milestone for our growing team. I am thrilled to announce that we have officially hired our very first apprentice. Bringing fresh talent into the fold is a responsibility I do not take lightly, and I am incredibly excited to mentor the next generation. How do you best support someone taking their very first steps into the professional world? We are focusing on building a strong foundation of daily habits and practical skills. I truly believe that small wins compound over time, creating massive growth for both the individual and the company. I cannot wait to watch this journey unfold. 🚀
Today, I’m excited to share that we’re hiring our first apprentice. This feels like a meaningful step for our team: investing in potential, creating a real pathway into the industry, and learning how to become better mentors along the way. Who could thrive with the right opportunity, patient guidance, and room to grow? We’re looking for curiosity, reliability, and a willingness to learn—not a perfect CV. We know small wins compound, and we’re ready to celebrate progress, build confidence, and give someone a strong start. If you know a motivated candidate, I’d love to connect. 🌱
anthropic/claude-sonnet-5: B — Both meet word count, 3 paragraphs, single question, phrase included, one emoji, no hashtags. Response 2 is slightly more engaging with a clear call to action, otherwise both are comparable in quality and clarity. / Both meet constraints (word count, structure, emoji, phrase). Response 1 is slightly more authentic and concise; Response 2 is a bit more clichéd with phrases like 'huge milestone' and 'massive growth'. | google/gemini-3.1-pro-preview: tie (flipped on swap) — Both responses followed all instructions and constraints perfectly. They both have the correct word count, paragraph structure, required phrase, and emoji placement. Both are well-written and suitable for LinkedIn. / Both responses followed all constraints perfectly. Response 1 is slightly better written and feels more natural as a LinkedIn post seeking candidates, whereas Response 2 announces an already completed hire but is also well-written. | x-ai/grok-4.5: tie — Both meet all constraints exactly (word counts 103/96, 3 paras, Q as para 2, phrase, 1 end emoji, no hashtags). Equally correct, clear, concise, and useful as LinkedIn posts. / Both meet all constraints exactly (word counts 96/103, 3 paras, Q as para 2, phrase present, 1 end emoji, no hashtags). Equally correct, clear, useful LinkedIn posts.
tieExact structureInstruction-following · 9.7 v 9.7Both follow exact structure, 3 sections, 4 items each, word counts within 3-8. Both correct, clear, concise. Response 2 has slightly less…
Produce a checklist for closing a small retail shop at night with EXACTLY this structure: a title line, then 3 sections headed 'Front of house', 'Back office', 'Security', each containing exactly 4 numbered items, each item 3-8 words. Nothing else before or after.
Nightly Retail Shop Closing Checklist Front of house 1. Wipe down all checkout counters. 2. Restock empty shelves and displays. 3. Sweep and mop the floors. 4. Turn off all display lights. Back office 1. Count cash in the register. 2. Lock the safe securely. 3. Log daily sales in binder. 4. Turn off office computers. Security 1. Check all windows are locked. 2. Set the main alarm system. 3. Lock the front entrance door. 4. Verify security cameras are recording.
Nightly Retail Shop Closing Checklist Front of house 1. Tidy displays and return misplaced items 2. Clean counters, floors, and fitting rooms 3. Restock bags, receipts, and supplies 4. Switch off music and lights Back office 1. Count cash and reconcile tills 2. Prepare bank deposit securely 3. File receipts and daily paperwork 4. Check inventory notes and deliveries Security 1. Lock stockroom and office doors 2. Set alarm system before leaving 3. Check all windows are secured 4. Lock entrance and exit doors
anthropic/claude-sonnet-5: tie — Both follow exact structure, 3 sections, 4 items each, word counts within 3-8. Both correct, clear, concise. Response 2 has slightly less redundancy (safe/cash count separate) but both equally valid overall. / Both follow exact structure, sections, 4 items each, word counts within 3-8, no extra text. Both correct and concise. Nearly identical quality; Response 2 slightly more consistent formatting with periods, Response 1 slightly less uniform but equally valid. | google/gemini-3.1-pro-preview: tie — Both responses followed all instructions and constraints perfectly, including the exact structure, section headers, item counts, and word count limits per item. Both provide highly useful and concise checklists for a retail shop. / Both responses followed all instructions and constraints perfectly, providing exactly the requested structure, section headers, item counts, and word counts per item. Both are highly useful and clear. | x-ai/grok-4.5: tie — Both match structure, word counts (3-8), headings, and item totals exactly; content is accurate, useful, clear. / Both match the exact required structure, section heads, 4 items each of 3-8 words, and nothing extra. Content is accurate, clear, and useful for the task.
tieQuick answer (short)Speed & cost · 10 v 10Both responses are identical, correct, concise, and fully meet the one-sentence requirement.
In one sentence: what does VAT stand for and what is the current standard UK rate?
VAT stands for Value Added Tax, and the current standard UK rate is 20%.
VAT stands for Value Added Tax, and the current standard UK rate is 20%.
anthropic/claude-sonnet-5: tie — Both responses are identical, correct, concise, and fully meet the one-sentence requirement. / Both responses are identical, accurate, concise, single-sentence answers that fully satisfy the task requirements. | google/gemini-3.1-pro-preview: tie — Both responses are identical, perfectly accurate, and follow all instructions, including the one-sentence constraint. / Both responses are identical, factually correct, and perfectly follow the constraint to provide the answer in exactly one sentence. | x-ai/grok-4.5: tie — Both identical, fully correct, one sentence as required, clear and useful. / Both identical, fully correct, one sentence as required, clear and useful.
tieInstant classificationSpeed & cost · 9.3 v 8.2Both correctly classify all items. Response 1 uses '/' separator matching task's own example format more closely; Response 2 uses '
Classify each as SPAM or NOT SPAM, output only two-column lines: 'You've won a £500 voucher click here' / 'Your Tesco order 8812 is ready for collection' / 'URGENT your account will be closed verify now' / 'Minutes from Thursday's board meeting attached'.
You've won a £500 voucher click here / SPAM Your Tesco order 8812 is ready for collection / NOT SPAM URGENT your account will be closed verify now / SPAM Minutes from Thursday's board meeting attached / NOT SPAM
SPAM | You've won a £500 voucher click here NOT SPAM | Your Tesco order 8812 is ready for collection SPAM | URGENT your account will be closed verify now NOT SPAM | Minutes from Thursday's board meeting attached
anthropic/claude-sonnet-5: tie (flipped on swap) — Both correctly classify all items. Response 1 uses '/' separator matching task's own example format more closely; Response 2 uses '|' and reverses column order, slight deviation from requested format. / Both correctly classify all four items and use consistent two-column format, differing only in column order, which wasn't specified explicitly. Both are clear and concise. | google/gemini-3.1-pro-preview: tie (flipped on swap) — Both responses correctly classify the messages and follow the strict formatting constraint to output only two-column lines. They just chose different separators and column orders, which is acceptable. / Both responses correctly classify the messages and follow the formatting constraints perfectly. Response 1 is chosen as the winner simply because placing the label in the first column is a more standard and readable convention for this type of classification task. | x-ai/grok-4.5: A — Both classify correctly. R1 matches exact two-column 'msg / LABEL' format required; R2 reverses columns and uses wrong separator, violating output constraint. / Both correctly classify all 4 (spam/not/spam/not) and output only the lines. Resp2 better matches the / separator shown in the task for two-column format.
tieShort translationSpeed & cost · 8.2 v 9.2Both accurate, informal translations. Response 1 provides helpful alternatives and nuance, useful for non-technical users. Response 2 is concise…
Translate to French, keeping the informal tone: 'Hey, we're running 10 minutes late - grab us a table if you get there first?'
Here is a natural, informal translation perfect for a text message: "Salut, on a 10 minutes de retard - prends-nous une table si tu arrives avant ?" *(Note: You can also use "Coucou" instead of "Salut" if you are very close with the person, and "garde-nous une table" instead of "prends-nous une table" which implies holding/saving the table).*
Salut, on a 10 minutes de retard — prends-nous une table si tu arrives avant nous ?
anthropic/claude-sonnet-5: tie (flipped on swap) — Both accurate, informal translations. Response 1 provides helpful alternatives and nuance, useful for non-technical users. Response 2 is concise and correct but slightly less helpful. Both follow instructions well. / Both translations are accurate and informal. Response 1 is concise and directly answers the task. Response 2 adds useful extra notes but is less concise, slightly less direct per conciseness criterion. | google/gemini-3.1-pro-preview: B — Both responses provide accurate, informal translations. Response 2 is better because it provides just the translation without unnecessary conversational filler, making it more concise and useful. / Response 1 provides the translation directly and concisely without unnecessary conversational filler, making it more useful for a busy professional. Both translations are accurate and maintain the requested informal tone. | x-ai/grok-4.5: tie (flipped on swap) — Both deliver accurate informal French translations. R1 adds helpful notes for a non-technical user; R2 is cleaner and slightly more precise. No constraints violated. / R1 is a direct, accurate informal translation matching the request exactly. R2 is also correct but adds unsolicited notes and framing, reducing concision and instruction adherence.
tieMedium generationSpeed & cost · 9 v 8.7Both are close to 150 words, cover all required elements clearly and warmly. Response 1 uses first-person which is slightly less consistent with…
Write a 150-word 'About us' section for a family-run garden centre in Yorkshire founded in 1987, now run by the founders' daughter, known for rare perennials and a small cafe.
Nestled in the heart of the beautiful North Yorkshire countryside, our garden centre has been a true labour of love. Founded by my parents in 1987 with a simple passion for growing, I am proud to continue our family legacy today as the second generation to tend these soils. Over the years, we have cultivated a reputation for offering an exceptional selection of rare and unusual perennials. Whether you are a seasoned horticulturist or a novice gardener, you will find hidden gems to bring colour and life to your borders. Beyond the plants, we believe a garden centre should be a place to linger. Our small, cosy cafe offers freshly baked cakes and warm Yorkshire brews, perfect for a mid-shop rest or a catch-up with old friends. We invite you to wander our peaceful paths, discover something truly extraordinary, and feel right at home in our little corner of Yorkshire.
Founded in 1987, our family-run garden centre has grown from a simple Yorkshire nursery into a welcoming destination for gardeners, plant lovers and neighbours alike. Today, it is proudly led by the founders’ daughter, who carries forward their practical knowledge, warm hospitality and belief that every garden can bring joy. We are especially known for our carefully chosen rare perennials: unusual, characterful plants selected for their beauty, resilience and ability to thrive in northern gardens. Whether you are creating a border, searching for something special, or simply seeking advice, our friendly team is here to help. After exploring the plants, pause in our small café for homemade cakes, light lunches and a proper cup of tea. From first-time growers to seasoned enthusiasts, we offer an unhurried place to discover, learn and leave inspired—rooted in Yorkshire, family values and a lifelong love of growing together, through every season of the year.
anthropic/claude-sonnet-5: tie — Both are close to 150 words, cover all required elements clearly and warmly. Response 1 uses first-person which is slightly less consistent with 'about us' business tone; Response 2 is more polished and consistent throughout. Both are strong, roughly equal quality. / Both are well-written, accurate to the brief, and cover founding, daughter, rare perennials, and cafe. Response 2 uses first-person voice, adding personal warmth. Neither strictly hits 150 words but both are close; no explicit banned words. Quality and usefulness are comparable. | google/gemini-3.1-pro-preview: A — Both responses hit exactly 150 words and include all required details. Response 1 is slightly better as it uses a first-person perspective ('my parents', 'I am proud'), which feels much more natural and engaging for a family-run business's 'About us' page compared to Response 2's third-person phrasing. / Both responses follow all instructions perfectly and hit the exact word count constraint. Response 2 is slightly better as it uses a first-person perspective ('my parents', 'I am proud'), which feels more authentic and personal for a family-run business 'About us' page. | x-ai/grok-4.5: tie — Both hit exactly 150 words and all required facts with clear, professional prose suited to the brief. Equally correct, compliant, useful and concise. / Both hit exactly 150 words and all required facts with clear, professional prose suited to the brief. Equally correct, compliant, useful and concise.
tieLong generationSpeed & cost · 6 v 6.2Both responses cover all required elements accurately and clearly. Response 1 hits ~400 words closely; Response 2 is longer with more detail but…
Write a detailed 400-word beginner's guide to setting up a home compost bin: location, what to add, what to avoid, how long it takes, and one common mistake.
Composting is a fantastic way to transform everyday kitchen scraps and yard waste into nutrient-rich soil for your garden. Setting up your first home compost bin is much easier than you might think. Here is a detailed beginner’s guide to get you started on your composting journey. **Location** Choosing the right spot is crucial for a healthy compost pile. Place your bin in a dry, shady area with good drainage and bare soil underneath, which allows beneficial microbes and worms to enter. You want it easily accessible year-round, ideally close to your kitchen door so you aren't deterred by bad weather when tossing in daily scraps. Direct sunlight can dry out the pile, while a completely hidden corner might make you forget to maintain it. **What to Add** A successful compost bin requires a balanced mix of "greens" and "browns." Greens provide nitrogen and include fruit and vegetable scraps, coffee grounds, tea bags, and fresh grass clippings. Browns provide carbon and include dry autumn leaves, shredded newspaper, plain cardboard, and small twigs. Aim for a ratio of roughly three parts browns to one part greens. Always chop larger items into smaller pieces to significantly speed up the decomposition process. **What to Avoid** Not everything belongs in your bin. Never add meat, bones, dairy products, or cooking oils, as these attract unwanted pests and create foul odors. Avoid diseased plants, pet waste, and weeds that have gone to seed, as an average backyard compost pile might not get hot enough to kill these pathogens. Also, skip glossy or coated paper, which contains harmful chemicals. **How Long It Takes** Patience is key in composting. Depending on the conditions, compost can take anywhere from three months to a full year to completely mature. Decomposition speeds up in warmer weather and when the pile is turned regularly to introduce oxygen. You will know your compost is ready when it is dark, crumbly, and smells like earthy forest soil. **One Common Mistake** The most common beginner mistake is letting the pile get too wet or too dry. A healthy compost pile should feel like a damp, wrung-out sponge. If it is too wet and slimy, it lacks oxygen and will smell; fix this by adding more browns and turning the pile. If it is too dry, decomposition stalls; fix this by adding water or fresh greens. By following these simple guidelines, you will soon produce a free, organic fertilizer that your plants will love, all while reducing your household waste.
# Beginner’s Guide to Setting Up a Home Compost Bin Composting turns kitchen scraps and yard waste into a dark, crumbly soil amendment that improves garden soil, helps plants retain moisture, and reduces household trash. A simple outdoor compost bin is an easy place to start. ## Choose the Right Location Place your compost bin outdoors on bare soil rather than concrete or a patio. Direct contact with soil allows helpful organisms, worms, and microbes to enter the pile. Choose a spot that is convenient enough to reach regularly, but not directly beside a door or outdoor seating area. A partially shaded location is ideal. Full sun can dry the compost too quickly, while deep shade may keep it overly wet and slow decomposition. Make sure you have access to a hose or water source, since compost should stay slightly damp. A bin can be a purchased plastic container, a wooden enclosure, a wire cylinder, or simply a covered pile. For most beginners, a bin with a lid helps keep out heavy rain and animals. ## What to Add Successful compost needs a mix of “greens” and “browns.” **Greens** are moist, nitrogen-rich materials that help feed decomposing organisms. Add items such as: - Fruit and vegetable scraps - Coffee grounds and paper filters - Tea leaves and uncoated tea bags - Fresh grass clippings - Crushed eggshells - Plant trimmings and wilted flowers **Browns** are dry, carbon-rich materials that create air pockets and prevent smells. Add: - Dry leaves - Shredded newspaper or plain cardboard - Small twigs - Straw - Wood chips or sawdust in small amounts - Paper towel tubes and uncoated paper towels Aim for roughly two to three parts browns for every one part greens. A good method is to cover each layer of food scraps with a layer of dry leaves, shredded cardboard, or newspaper. ## What to Avoid Do not add meat, fish, bones, dairy products, oils, grease, or cooked foods with sauces. These can attract rats, raccoons, flies, and other pests. Avoid pet waste, especially from cats and dogs, because it may contain harmful pathogens. Also skip diseased plants, weeds that have gone to seed, glossy paper, treated wood, and coal or charcoal ash. ## How Long Compost Takes Compost can take anywhere from two months to a year. A well-managed pile that is turned regularly, kept moist, and filled with small pieces may produce usable compost in two to four months. A pile left mostly alone may take six to twelve months. Finished compost is dark brown, earthy-smelling, and crumbly. You should not be able to recognize most of the original materials. ## One Common Mistake A common beginner mistake is adding too many wet food scraps without enough browns. This creates a soggy, smelly pile. If your compost smells rotten, mix in dry leaves, shredded cardboard, or newspaper and turn the pile gently to add air. The correct balance should feel like a wrung-out sponge: moist, but not dripping wet.
anthropic/claude-sonnet-5: tie (flipped on swap) — Both responses cover all required elements accurately and clearly. Response 1 hits ~400 words closely; Response 2 is longer with more detail but still useful and well-organized. Both equally satisfy the task, with slight length overage in Response 2, but no explicit constraint violation to check against. / Both cover all required points accurately and clearly. Response 1 is slightly more detailed and organized with lists, aiding a non-technical reader. Response 2 is solid but slightly less thorough. Neither has an explicit word-count constraint violation apparent, though 'detailed 400-word' guideline is approximate for both. | google/gemini-3.1-pro-preview: tie — Both responses fail the strict 400-word count constraint. However, Response 2 is slightly better structured for a busy professional, utilizing bullet points for the 'what to add' section, making it easier to read and reference quickly. Both are factually accurate and cover all required topics. / Both responses fail to meet the exact 400-word count constraint. Response 1 is better formatted and easier to read for a beginner, utilizing bullet points and clear headings. | x-ai/grok-4.5: tie — Both cover all required topics accurately and usefully with clear structure. R2 superior in lists, detail, and scannability for busy readers. Neither hits 400 words (R1 ~489, R2 ~520). / Both exceed 400 words (R1~510, R2~460) and cover all topics accurately/clearly. R1 more detailed/useful lists and structure for beginners.
What they cost
- api price in per 1m$0.42OpenRouter API12 Aug 2026stale
- api price in per 1m$0.42OpenRouter API12 Aug 2026stale
- api price in per 1m$0.42OpenRouter API12 Aug 2026stale
- api price in per 1m$0.42OpenRouter API12 Aug 2026stale
- api price out per 1m$1.4OpenRouter API12 Aug 2026stale
- api price in per 1m$1OpenRouter API12 Aug 2026stale
- api price out per 1m$6OpenRouter API12 Aug 2026stale
Speed, measured
Median response time across all 30 tasks: 12.8s (GLM 5.2) vs 2.7s (GPT-5.6 Terra).
Receipts — every citation, raw outputs
The full run — prompts, both raw outputs per task, judge reasoning, latencies, token counts and per-call cost — is downloadable as JSON. The task suite is versioned (suite-2026-07) and published on the methodology page.
- judge panelanthropic/claude-sonnet-5, google/gemini-3.1-pro-preview, x-ai/grok-4.5our run (raw outputs)12 Aug 2026verified
- judge swap agreement0.533our run (raw outputs)12 Aug 2026verified
- judge swap kappa0.228our run (raw outputs)12 Aug 2026verified
- median latency ms a12762our run (raw outputs)12 Aug 2026verified
- median latency ms b2735our run (raw outputs)12 Aug 2026verified
- panel swap flip rate0.367our run (raw outputs)12 Aug 2026verified
- panel unanimous rate0.433our run (raw outputs)12 Aug 2026verified
- run cost a usd0.1015our run (raw outputs)12 Aug 2026verified
- run cost b usd0.0462our run (raw outputs)12 Aug 2026verified
- score a3our run (raw outputs)12 Aug 2026verified
- score b2our run (raw outputs)12 Aug 2026verified
- suite winswriting: a 1/b 0/tie 4 · coding: a 0/b 0/tie 5 · reasoning: a 0/b 0/tie 5 · extraction: a 2/b 1/tie 2 · instruction: a 0/b 1/tie 4 · speed cost: a 0/b 0/tie 5our run (raw outputs)12 Aug 2026verified
- tasks total30our run (raw outputs)12 Aug 2026verified
- ties25our run (raw outputs)12 Aug 2026verified
a narrow win on the tasks that separated them (5 of 30 tasks were decisive) — close enough that the loser is still worth a look.
“[Critical] GLM-5.2 API is unusable due to severe rate limiting — 2 consecutive days of near-total outage”on GLM 5.2
“Regression: encrypted MultiAgentV2 messages remove readable task audit trail”on GPT-5.6 Terra