Battles / Best-value API
Dead heat, 4–4
level across 30 tasks — genuinely nothing between them.
GPT-5.6 Terra vs Grok 4.5 · Best-value APImodels in this battle: GPT-5.6 Terra vs Grok 4.5
Show the full workings
The evidence
Suite-by-suite
blue = GPT-5.6 Terra wins · grey = ties · white = Grok 4.5 wins (5 tasks per suite)
Round-by-round — all 30 tasks
tieCold emailWriting · 8.5 v 8.8Both meet word limit, friendly professional tone, clear CTA, no buzzwords. Response 1 mentions specific brewery-relevant tasks adding usefulness…
Write a cold email (max 120 words) from the owner of a 6-person UK accounting firm to a local brewery, offering a free first-month bookkeeping trial. Friendly but professional, one clear call to action, no buzzwords.
Subject: A free month of bookkeeping for your brewery Hi [Name], I’m [Your Name], owner of [Firm Name], a small UK accounting firm based in [Town]. We work with local businesses that need clear, reliable bookkeeping without adding to their workload. For breweries, that can mean keeping on top of invoices, supplier costs, payroll and VAT while you focus on making and selling great beer. To show how we work, I’d like to offer you your first month of bookkeeping free, with no obligation to continue afterwards. Would you be open to a 15-minute call next week to see if it would be useful? Best regards, [Your Name] [Firm Name] [Phone number]
**Subject:** Free first-month bookkeeping for [Brewery Name] Hi [Name], I’m [Your Name], owner of [Firm Name], a six-person accounting firm based nearby. We handle the books for a handful of local businesses and thought your brewery might find a bit of extra support useful. I’d like to offer you a free first month of bookkeeping—no cost, no commitment. We’ll tidy the records, handle the day-to-day entries, and show you exactly how we work. If that sounds helpful, just reply to this email and we’ll get started. All the best, [Your Name] Owner, [Firm Name] [Phone] | [Website] (Word count: 98)
anthropic/claude-sonnet-5: tie (flipped on swap) — Both meet word limit, friendly professional tone, clear CTA, no buzzwords. Response 1 mentions specific brewery-relevant tasks adding usefulness; Response 2 explicitly states word count, mentions team size per no requirement but fine. Both very similar quality. / Both meet word limit, tone, and CTA requirements. Response 2 has a stronger, more specific CTA (15-minute call) and brewery-relevant details, making it slightly more useful and concrete than Response 1's generic reply prompt. | google/gemini-3.1-pro-preview: B — Both responses follow the constraints well and are under the 120-word limit. Response 2 explicitly includes the '6-person' firm detail from the prompt, whereas Response 1 only mentions being a 'small' firm. Response 2 is slightly more concise and direct. / Response 1 includes the specific detail of being a 6-person firm as requested, whereas Response 2 just says 'small'. Both follow the word count constraint and provide a clear, professional email with a single call to action. | deepseek/deepseek-v4-pro: tie (flipped on swap) — Both emails meet all constraints: under 120 words, no buzzwords, one clear CTA, friendly yet professional. Response 1 tailors to breweries; Response 2 is more direct. Equally effective. / Response 1 is more concise, has a lower-friction CTA (just reply), and precisely states 'six-person firm' as required. Response 2 is slightly wordier and its CTA (15-minute call) is a bigger ask. Both are accurate, friendly, and buzzword-free.
Grok 4.5Product descriptionWriting · 5 v 9.5R1 is 59 words, missing the 60-80 word requirement; R2 hits 70 words correctly. Both avoid cliches and cover features accurately, but R2 fully…
Write a 60-80 word product description for a stainless-steel insulated water bottle (750ml, keeps drinks cold 24h/hot 12h, leakproof, fits car cup holders). Target audience: gym-goers. Avoid cliches like 'stay hydrated in style'.
Built for demanding workouts and busy schedules, this 750ml stainless-steel insulated bottle keeps water cold for up to 24 hours and coffee or tea hot for up to 12. Its leakproof lid protects your gym bag, while the slim shape fits standard car cup holders for easy post-training commutes. Durable, reusable, and easy to carry from warm-up to recovery.
This 750ml stainless-steel insulated water bottle is built for gym enthusiasts who demand reliable performance. It keeps your water icy cold for 24 hours or your post-workout tea hot for 12, ensuring refreshment throughout intense sessions. The leakproof design means no spills in your gym bag, while its slim shape fits perfectly in car cup holders for easy transport. Durable and practical, it's the ideal companion for your fitness routine.
anthropic/claude-sonnet-5: B — R1 is 59 words, missing the 60-80 word requirement; R2 hits 70 words correctly. Both avoid cliches and cover features accurately, but R2 fully complies and is equally clear and useful. / Both are accurate, clear, and avoid cliches. R1 meets the 60-80 word count (70 words), while R2 falls short at 59 words, violating the explicit word count constraint. R1 is slightly more detailed and complete. | google/gemini-3.1-pro-preview: B — Response 1 falls just short of the 60-word minimum constraint (59 words). Response 2 meets the word count constraint (70 words), includes all required features, targets the correct audience, and avoids the banned cliché. / Response 1 followed all instructions, including the 60-80 word count constraint (70 words). Response 2 missed the minimum word count constraint (59 words). Both effectively targeted gym-goers and included all required features without using cliches. | deepseek/deepseek-v4-pro: B — Response 1 word count 59 (below required 60-80), violates constraint. Response 2 meets all criteria, accurate, avoids cliché, targets gym-goers, clear. Winner: 2. / Response 1 meets all constraints, accurate, clear, useful. Response 2 is under word count (59 words), violating the 60-80 word requirement. Both avoid cliches and target gym-goers well. Winner: 1.
tieSummarise messy notesWriting · 9.2 v 9.5Both are accurate, concise 5-bullet summaries covering all key points within instructions. Response 1 slightly more polished formatting, Response…
Turn these messy meeting notes into a clean 5-bullet summary a director could read in 20 seconds: 'ok so launch slipped again, marc says supplier issue?? maybe 3 wks. sarah wants to push the email campaign anyway but tom thinks wait. budget - we're 8k over, mostly the packaging redo. Q: do we tell retail partners now or after new date confirmed. also NEED to hire the warehouse temp before august rush. next mtg tues.'
- **Launch delayed again:** Supplier issue is expected to push timing by approximately **three weeks**; revised date is not yet confirmed. - **Marketing decision pending:** Sarah recommends proceeding with the email campaign, while Tom advises waiting until the new launch date is set. - **Budget overrun:** Project is currently **$8K over budget**, driven primarily by the packaging redesign. - **Retail communications:** Decide whether to notify retail partners now or wait until the revised launch date is confirmed. - **Staffing and next steps:** Hire a warehouse temporary worker before the August rush; next meeting is **Tuesday**.
- **Launch delay:** Slipped again (supplier issue per Marc); potential 3-week push. - **Email campaign:** Sarah wants to proceed now; Tom recommends waiting for new launch date. - **Budget:** $8k over, driven mainly by packaging redo. - **Retail partners:** Decide whether to notify now or only after revised date is firm. - **Staffing + next steps:** Hire warehouse temp before August rush; next meeting Tuesday.
anthropic/claude-sonnet-5: tie — Both are accurate, concise 5-bullet summaries covering all key points within instructions. Response 1 slightly more polished formatting, Response 2 slightly more concise. Equally usable for a director. / Both accurately capture all key points in 5 clear bullets, meeting the 20-second-read goal. Response 2 is slightly more polished with bold emphasis but marginally longer/denser. Response 1 is more concise. Both fully satisfy the task's implicit constraints. | google/gemini-3.1-pro-preview: B — Both responses accurately capture the meeting notes in exactly 5 bullets. Response 2 is slightly punchier and more concise, making it slightly better suited for a quick 20-second read by a busy director. / Both responses accurately summarize the notes into exactly 5 bullets. Response 1 is slightly more concise and punchy, making it better suited for a busy director to read in under 20 seconds. | deepseek/deepseek-v4-pro: tie — Both accurately capture all key points in exactly 5 clean, scannable bullets. Response 1 is slightly more formal, Response 2 more concise; both are highly useful and clear. / Both accurately capture all facts, exactly follow the 5-bullet instruction, are clear and concise, and useful for a director. No constraints violated.
Grok 4.5Tone rewriteWriting · 8.8 v 9.5Both responses are warm, accountable, factually consistent, under 90 words, and clear. Response 2 slightly more explicit in taking responsibility…
Rewrite this complaint reply so it is warm, takes responsibility, and keeps the same facts, in under 90 words: 'Your order was delayed because of a system error on our side. The refund has been processed and will arrive in 3-5 days. We cannot offer further compensation. Let us know if you have questions.'
We’re very sorry for the delay with your order. A system error on our side caused the issue, and we understand how frustrating this must be. Your refund has now been processed and should arrive within 3–5 days. While we’re unable to offer additional compensation, we appreciate your patience and are here to help with any questions.
We’re truly sorry your order was delayed due to a system error on our side—that’s on us, and we take full responsibility. Your refund has already been processed and should arrive in 3–5 days. We’re unable to offer further compensation, but please don’t hesitate to reach out if you have any questions. We’re here to help.
anthropic/claude-sonnet-5: tie — Both responses are warm, accountable, factually consistent, under 90 words, and clear. Response 2 slightly more explicit in taking responsibility, but both are essentially equal in quality and usefulness. / Both responses are warm, accurate, under 90 words, and retain all facts. R2 shows slightly more empathy by acknowledging frustration, but both are equally clear and concise, making them essentially tied in quality. | google/gemini-3.1-pro-preview: B — Both responses follow all instructions and constraints perfectly. Response 2 is slightly better as it explicitly takes responsibility ('that's on us, and we take full responsibility'), making it feel warmer and more accountable. / Both responses follow all instructions and constraints perfectly. Response 1 is slightly better because it explicitly states 'we take full responsibility', which directly addresses the prompt's requirement to take responsibility in a very clear and warm manner. | deepseek/deepseek-v4-pro: B — Response 2 takes explicit responsibility ('we take full responsibility') and is warmer, making it more useful. Response 1 is empathetic but doesn't enhance responsibility. Both are factually accurate and under word limit. / Response 1 explicitly takes responsibility ('that’s on us') and is warmer, while both keep facts, are under 90 words, and are clear. Response 2 is empathetic but less direct in owning the error.
tieHeadlinesWriting · 9.3 v 9.3Both meet word limits and category mix. Response 2 explicitly labels categories, aiding clarity and usefulness for a non-technical reader, while…
Write 5 headline options (max 9 words each) for a blog post about how small restaurants can use AI to reduce food waste. Mix: 2 practical, 2 curiosity-driven, 1 with a number.
- Cut Food Waste With AI-Powered Inventory Forecasting - Use AI to Turn Leftovers Into Profits - The AI Secret Small Restaurants Are Missing - Can AI Predict Tomorrow’s Food Waste? - 5 Ways AI Helps Restaurants Waste Less Food
**Headline options:** 1. **Practical:** Use AI Tools to Cut Restaurant Food Waste 2. **Practical:** How Small Restaurants Track Inventory with AI 3. **Curiosity-driven:** What If Your Kitchen Predicted Tomorrow’s Waste? 4. **Curiosity-driven:** The Quiet AI Fix Hiding in Your Freezer 5. **With a number:** 7 AI Ways Small Restaurants Slash Food Waste
anthropic/claude-sonnet-5: tie (flipped on swap) — Both meet word limits and category mix. Response 2 explicitly labels categories, aiding clarity and usefulness for a non-technical reader, while Response 1 is slightly less transparent about which is which. / Both meet the 5-headline mix requirement, word counts under 9, and categories (2 practical, 2 curiosity, 1 number). Response 1 labels categories explicitly, aiding clarity; Response 2 is more concise and punchy. Both are correct and useful; quality is comparable. | google/gemini-3.1-pro-preview: tie (flipped on swap) — Both responses followed all instructions perfectly, including word counts and the specific mix of headline types. Response 1 is slightly better as it presents the headlines cleanly without unnecessary labels, making it more concise and ready to use. / Both responses followed all instructions perfectly, including the word count and specific mix of headline types. Response 1 is slightly better because it explicitly labels the categories, making it easier for the user to see that the constraints were met. | deepseek/deepseek-v4-pro: tie — Both responses provide 5 headlines under 9 words with the required mix (2 practical, 2 curiosity-driven, 1 numbered). They are accurate, clear, and useful for a non-technical audience. / Both responses contain 5 headlines, each ≤9 words, with 2 practical, 2 curiosity-driven, and 1 numbered. Factual and useful, with no errors or breaches.
tieBug fixCoding · 9.7 v 9.5Both correct and functionally identical, but Response 1 keeps explanation inline as one line per instructions, while Response 2 uses markdown code…
This JavaScript function should return the median of a numeric array but gives wrong answers for even-length arrays and mutates the input. Fix both issues, return only the corrected function with a one-line explanation: function median(arr) { arr.sort(); const mid = Math.floor(arr.length / 2); return arr[mid]; }
function median(arr) { const sorted = [...arr].sort((a, b) => a - b); const mid = Math.floor(sorted.length / 2); return sorted.length % 2 ? sorted[mid] : (sorted[mid - 1] + sorted[mid]) / 2; } // Copies before numeric sorting and averages the two middle values for even-length arrays.
```js function median(arr) { const sorted = [...arr].sort((a, b) => a - b); const mid = Math.floor(sorted.length / 2); return sorted.length % 2 ? sorted[mid] : (sorted[mid - 1] + sorted[mid]) / 2; } ``` Copies input (no mutation), numeric sort, averages two middle values for even length.
anthropic/claude-sonnet-5: tie (flipped on swap) — Both correct and functionally identical, but Response 1 keeps explanation inline as one line per instructions, while Response 2 uses markdown code block and separate explanation line, slightly deviating from 'return only the corrected function with a one-line explanation' format. / Both responses provide identical correct code with proper fixes and a concise one-line explanation, differing only in code block formatting. | google/gemini-3.1-pro-preview: tie (flipped on swap) — Both responses correctly fix the mutation, numeric sorting, and even-length array issues. Response 1 is slightly better as it strictly returns only the code and the explanation as a comment, avoiding markdown formatting which better adheres to the 'return only' constraint. / Both responses correctly fix the mutation, numeric sorting, and even-length array issues. Response 1 is slightly better formatted with markdown code blocks, making it easier to read. | deepseek/deepseek-v4-pro: tie — Both responses fix mutation and even-length median, add numeric sort. No constraint violations. Identical quality. / Both fix mutation and even-length median correctly, provide a concise one-line explanation, and are equally useful.
GPT-5.6 TerraSmall featureCoding · 9.5 v 8.8Both correctly implement chunking with oversized invoice handling. Response 1 is slightly more explicit about oversized invoice logic and adds an…
Write a Python function `chunk_invoices(invoices, max_total)` that takes a list of dicts like {'id': 'A1', 'amount': 120.5} and groups them into batches where each batch's summed amount does not exceed max_total. A single invoice larger than max_total goes in its own batch. Preserve order. Include 3 test cases using assert.
```python def chunk_invoices(invoices, max_total): """ Group invoices into ordered batches whose total amount does not exceed max_total. Any single invoice larger than max_total is placed in its own batch. Args: invoices: List of dictionaries containing at least an "amount" key. max_total: Maximum allowed total per batch. Returns: A list of invoice batches (each batch is a list of invoice dicts). """ if max_total <= 0: raise ValueError("max_total must be greater than 0") batches = [] current_batch = [] current_total = 0 for invoice in invoices: amount = invoice["amount"] # Oversized invoices always get their own batch. if amount > max_total: if current_batch: batches.append(current_batch) current_batch = [] current_total = 0 batches.append([invoice]) continue # Start a new batch if adding this invoice would exceed the limit. if current_batch and current_total + amount > max_total: batches.append(current_batch) current_batch = [] current_total = 0 current_batch.append(invoice) current_total += amount if current_batch: batches.append(current_batch) return batches # Test 1: Normal batching while preserving order invoices = [ {"id": "A1", "amount": 120.5}, {"id": "A2", "amount": 50}, {"id": "A3", "amount": 80}, {"id": "A4", "amount": 20}, ] assert chunk_invoices(invoices, 200) == [ [{"id": "A1", "amount": 120.5}, {"id": "A2", "amount": 50}], [{"id": "A3", "amount": 80}, {"id": "A4", "amount": 20}], ] # Test 2: An oversized invoice gets its own batch invoices = [ {"id": "A1", "amount": 40}, {"id": "A2", "amount": 250}, {"id": "A3", "amount": 60}, ] assert chunk_invoices(invoices, 100) == [ [{"id": "A1", "amount": 40}], [{"id": "A2", "amount": 250}], [{"id": "A3", "amount": 60}], ] # Test 3: Empty input assert chunk_invoices([], 100) == [] ```
```python def chunk_invoices(invoices, max_total): """Group invoices into batches where each batch sum <= max_total. Invoices larger than max_total go in their own batch. Order is preserved. """ batches = [] current_batch = [] current_sum = 0.0 for inv in invoices: amount = inv["amount"] if current_batch and current_sum + amount > max_total: batches.append(current_batch) current_batch = [] current_sum = 0.0 current_batch.append(inv) current_sum += amount if current_batch: batches.append(current_batch) return batches # Test cases inv1 = [ {"id": "A", "amount": 10.0}, {"id": "B", "amount": 20.0}, {"id": "C", "amount": 30.0}, ] assert chunk_invoices(inv1, 40.0) == [ [{"id": "A", "amount": 10.0}, {"id": "B", "amount": 20.0}], [{"id": "C", "amount": 30.0}], ] inv2 = [{"id": "X", "amount": 100.0}] assert chunk_invoices(inv2, 50.0) == [[{"id": "X", "amount": 100.0}]] inv3 = [ {"id": "A", "amount": 10.0}, {"id": "B", "amount": 100.0}, {"id": "C", "amount": 10.0}, ] assert chunk_invoices(inv3, 50.0) == [ [{"id": "A", "amount": 10.0}], [{"id": "B", "amount": 100.0}], [{"id": "C", "amount": 10.0}], ] print("All tests passed.") ```
anthropic/claude-sonnet-5: A — Both correctly implement chunking with oversized invoice handling. Response 1 is slightly more explicit about oversized invoice logic and adds an empty-input test, giving marginally better coverage and robustness (max_total validation). / Both are correct and follow instructions. Response 2 explicitly handles oversized invoices in a clearer dedicated branch, includes an empty input edge case, and has slightly more robust validation, making it marginally more thorough and clear. | google/gemini-3.1-pro-preview: tie (flipped on swap) — Both responses provide correct, well-written Python functions that meet all constraints and handle oversized invoices properly. Response 2 has slightly more elegant logic, while Response 1 is more explicit, but both are excellent. / Both responses provide correct and well-tested solutions. Response 1 is slightly preferred for its more concise and elegant logic that naturally handles oversized invoices without needing a special conditional branch. | deepseek/deepseek-v4-pro: A — Response 1 explicitly handles oversized invoices and invalid input, making it more robust and clear. Response 2’s batch splitting works but is less explicit. Both satisfy all constraints and include 3 assert tests. / Both correctly batch invoices and include 3 asserts. Response 2 adds an explicit oversized-invoice branch, validation for max_total<=0, and more detailed docstring, making it slightly more robust and useful for a professional.
tieExplain codeCoding · 8.3 v 8.8Both are concise and under 150 words. Response 2 correctly identifies the more critical bug: the shared global cache causes key collisions across…
Explain to a junior developer, in under 150 words, what this code does and one risk of using it: const cache = new Map(); function memo(fn) { return (...args) => { const k = JSON.stringify(args); if (!cache.has(k)) cache.set(k, fn(...args)); return cache.get(k); }; }
This code creates a simple memoization wrapper: it remembers a function’s previous results. - `cache` stores results in a `Map`. - `memo(fn)` returns a new function. - When called, it converts the arguments into a JSON string key. - If that key is not cached, it runs `fn(...args)` and saves the result. - Later calls with the same serialized arguments return the saved result instead of rerunning `fn`. Example: wrapping an expensive calculation can make repeated calls faster. One risk: the cache is global and never removes entries, so it can grow forever and cause memory issues. Also, `JSON.stringify` may not reliably distinguish all argument types (for example, functions, `undefined`, or circular objects).
**What it does:** This is a simple memoization helper. `memo` takes any function `fn` and returns a wrapped version. On each call it turns the arguments into a string key via `JSON.stringify`, checks a shared `Map` cache, and either returns the stored result or calls `fn`, stores the result, then returns it. Repeated calls with the same arguments skip re-execution. **One risk:** The cache is module-level and never cleared, so it grows forever (memory leak). It is also shared by every function you wrap, so different functions can collide on the same key.
anthropic/claude-sonnet-5: tie (flipped on swap) — Both are concise and under 150 words. Response 2 correctly identifies the more critical bug: the shared global cache causes key collisions across different wrapped functions, a subtle but important risk that Response 1 misses. / Both are accurate and concise, under 150 words. Response 2 adds useful nuance about JSON.stringify limitations and includes an example, slightly more helpful for a junior dev without exceeding length constraints. | google/gemini-3.1-pro-preview: B — Both responses meet the word count constraint and explain the code well. Response 2 is superior because it correctly identifies a critical bug in the provided snippet: the cache is shared across all wrapped functions, leading to key collisions. / Both responses meet all constraints and explain the code well. Response 1 is slightly better because it correctly identifies a critical bug in the provided code: the cache is shared across all wrapped functions, leading to key collisions. | deepseek/deepseek-v4-pro: tie (flipped on swap) — Response 1 explains more clearly with bullet points and an example, better for a junior developer. Both mention more than one risk against the 'one risk' instruction, but that affects both equally. Response 1's extra risk is still a minor deviation, yet overall more informative. / Both responses correctly explain memoization and a risk, are under 150 words, and are clear. Response 1 notes cache collision; Response 2 notes JSON.stringify limitations. Equally useful.
tieSQL queryCoding · 9.7 v 9.7Both queries are correct, use window functions, group by customer id to avoid name collisions, and filter 2025 dates properly. Nearly identical…
Given tables orders(id, customer_id, created_at, total) and customers(id, name, country), write a single SQL query returning each country's top 3 customers by lifetime spend in 2025, with columns country, name, total_spend, rank. Use a window function. Standard PostgreSQL.
```sql WITH customer_spend AS ( SELECT c.country, c.name, c.id AS customer_id, SUM(o.total) AS total_spend FROM customers c JOIN orders o ON o.customer_id = c.id WHERE o.created_at >= DATE '2025-01-01' AND o.created_at < DATE '2026-01-01' GROUP BY c.country, c.name, c.id ), ranked_customers AS ( SELECT country, name, total_spend, RANK() OVER ( PARTITION BY country ORDER BY total_spend DESC ) AS rank FROM customer_spend ) SELECT country, name, total_spend, rank FROM ranked_customers WHERE rank <= 3 ORDER BY country, rank, name; ```
```sql WITH customer_spend AS ( SELECT c.country, c.name, SUM(o.total) AS total_spend FROM customers c JOIN orders o ON o.customer_id = c.id WHERE o.created_at >= '2025-01-01' AND o.created_at < '2026-01-01' GROUP BY c.id, c.country, c.name ), ranked AS ( SELECT country, name, total_spend, RANK() OVER (PARTITION BY country ORDER BY total_spend DESC) AS rank FROM customer_spend ) SELECT country, name, total_spend, rank FROM ranked WHERE rank <= 3 ORDER BY country, rank; ```
anthropic/claude-sonnet-5: tie — Both queries are correct, use window functions, group by customer id to avoid name collisions, and filter 2025 dates properly. Nearly identical logic and clarity; Response 1 adds secondary name ordering, a minor plus but not required. / Both queries are correct, use window functions properly, follow instructions, and are clear. Nearly identical logic; response 2 adds minor tie-breaking with name, a slight but negligible improvement. | google/gemini-3.1-pro-preview: tie — Both responses provide accurate, well-structured PostgreSQL queries that perfectly meet all requirements, including the use of window functions and correct column outputs. They both correctly group by customer ID to handle potential duplicate names. / Both responses provide accurate, well-structured PostgreSQL queries that perfectly meet all requirements, including the use of window functions and the correct output columns. They both correctly filter for the year 2025. | deepseek/deepseek-v4-pro: tie — Both queries are accurate, use RANK() window function, filter 2025 correctly, and return required columns. Response 1 explicitly casts dates and orders by name for tie-breaking; Response 2 is concise. Both meet all instructions perfectly. No constraints violated. / Both queries correctly compute top 3 per country by 2025 spend with window function. They are accurate, meet all instructions, and are clear. Response 2 adds deterministic tiebreaker ordering, but neither deviates from requirements.
tieRegexCoding · 9 v 7.7Response 1's regex is more accurate to official UK postcode format rules, handles GIR 0AA, and explicitly notes a valid edge case (BFPO). Response…
Write a regex that matches UK postcodes like 'SW1A 1AA', 'M1 1AE', 'B33 8TH' (allow lowercase and optional space), and a one-line JavaScript example using it to validate a form field. Briefly note one edge case your regex does NOT handle.
```regex /^(?:GIR ?0AA|[A-PR-UWYZ](?:\d{1,2}|[A-HK-Y]\d[\dABEHMNPRVWXY]?) ?\d[ABD-HJLNP-UW-Z]{2})$/i ``` ```js const isValidPostcode = /^(?:GIR ?0AA|[A-PR-UWYZ](?:\d{1,2}|[A-HK-Y]\d[\dABEHMNPRVWXY]?) ?\d[ABD-HJLNP-UW-Z]{2})$/i.test(document.querySelector('#postcode').value.trim()); ``` This accepts lowercase and an optional space, e.g. `sw1a1aa`, `M1 1AE`, and `B338TH`. Edge case: it does not handle BFPO military addresses such as `BFPO 801`, which are not standard UK postcode formats.
**Regex:** `/^[A-Z]{1,2}\d[A-Z\d]?\s*\d[A-Z]{2}$/i` **One-line JS example:** `const valid = /^[A-Z]{1,2}\d[A-Z\d]?\s*\d[A-Z]{2}$/i.test(document.getElementById('postcode').value);` **Edge case not handled:** Special postcodes such as `GIR 0AA` (and similar non-standard formats).
anthropic/claude-sonnet-5: A — Response 1's regex is more accurate to official UK postcode format rules, handles GIR 0AA, and explicitly notes a valid edge case (BFPO). Response 2's regex is looser and its noted edge case is actually handled by their claim inaccurately (GIR 0AA not covered by their regex but not mentioned as such). / Response 2's regex is more accurate to real UK postcode rules and includes GIR 0AA handling, plus a valid edge case note. Response 1's regex is simpler but less precise. Both meet format requirements. | google/gemini-3.1-pro-preview: tie (flipped on swap) — Response 1 provides a highly accurate, standard regex for UK postcodes, whereas Response 2 provides a naive approximation that allows many invalid formats. Both follow all instructions, but Response 1 is much more useful and correct. / Both responses follow all instructions perfectly. Response 1 provides a simpler regex which is easier to read, while Response 2 provides a highly accurate, production-ready regex. Both correctly include the JS one-liner and an edge case. | deepseek/deepseek-v4-pro: tie (flipped on swap) — Response 1's regex precisely follows official UK postcode rules and excludes invalid characters; Response 2 is too permissive, accepting many non-conforming strings. Both satisfy task instructions, but 1 is far more accurate and professionally useful. / Response 1 is simpler, more concise, and equally correct for the given examples, making it more useful for a non-technical professional. Response 2 is more comprehensive but less clear.
GPT-5.6 TerraMulti-step logicReasoning · 9.7 v 9Both compute correctly (£60 vs £146), with response 2 having a minor labour miscalc (528 instead of 468+60=528, actually correct). Response 1 is…
A bakery sells loaves at £3.20. Ingredients cost £1.10/loaf, labour £0.90/loaf, fixed costs £480/week. They sell 450 loaves/week. A supplier offers ingredients at £0.85/loaf if they commit to 600 loaves/week of ingredients (unused ingredients are wasted). Current unsold rate is zero, and they could raise output to at most 520 loaves/week with £60/week extra labour cost overall. Should they take the deal? Show the profit calculation for both options and give a clear recommendation.
### Option 1: Keep current arrangement (450 loaves/week) - Revenue: \(450 \times £3.20 = £1,440\) - Ingredients: \(450 \times £1.10 = £495\) - Labour: \(450 \times £0.90 = £405\) - Fixed costs: £480 \[ \text{Profit} = 1,440 - 495 - 405 - 480 = \boxed{£60\text{/week}} \] --- ### Option 2: Take supplier deal and raise production to 520 loaves/week They must buy ingredients for 600 loaves, even though they can only make 520. Thus 80 loaves' worth of ingredients are wasted. - Revenue: \(520 \times £3.20 = £1,664\) - Ingredients purchased: \(600 \times £0.85 = £510\) - Labour: \(520 \times £0.90 = £468\) - Extra labour cost: £60 - Fixed costs: £480 \[ \text{Profit} = 1,664 - 510 - 468 - 60 - 480 \] \[ \text{Profit} = \boxed{£146\text{/week}} \] --- ### Comparison | Option | Weekly profit | |---|---:| | Stay at 450 loaves | £60 | | Take deal and make/sell 520 loaves | £146 | \[ £146 - £60 = \boxed{£86\text{ more profit per week}} \] ### Recommendation **Take the deal only if the bakery can sell the increased 520-loaf output.** On that basis, it raises weekly profit by **£86**, despite wasting ingredients for 80 loaves. However, if sales remain limited to 450 loaves, the deal is unattractive: \[ \text{Profit at 450 with deal} = 1,440 - 510 - 405 - 480 = \boxed{£45} \] So if the bakery cannot reliably sell more than 450 loaves, it should **not** take the deal.
The current weekly profit is calculated as follows: - Revenue: \(450 \times 3.20 = 1440\) - Ingredients: \(450 \times 1.10 = 495\) - Labour: \(450 \times 0.90 = 405\) - Fixed costs: \(480\) - Total costs: \(495 + 405 + 480 = 1380\) - Profit: \(1440 - 1380 = 60\) Under the supplier deal the bakery must purchase ingredients for 600 loaves (cost \(600 \times 0.85 = 510\)) while production (and sales) can be raised to at most 520 loaves at an extra labour cost of \(60\). The resulting weekly profit is: - Revenue: \(520 \times 3.20 = 1664\) - Ingredients: \(510\) - Labour: \(520 \times 0.90 + 60 = 528\) - Fixed costs: \(480\) - Total costs: \(510 + 528 + 480 = 1518\) - Profit: \(1664 - 1518 = 146\) (If the bakery took the deal but kept output at 450 loaves the profit would fall to 45, so the capacity increase is essential.) The deal raises weekly profit from \(60\) to \(146\). The bakery should therefore accept it. Final answer: yes (profit £60 → £146)
anthropic/claude-sonnet-5: A — Both compute correctly (£60 vs £146), with response 2 having a minor labour miscalc (528 instead of 468+60=528, actually correct). Response 1 is more detailed, clear tables, caveats on sales feasibility, slightly more thorough recommendation. / Both correctly compute £60 vs £146 profit and recommend the deal contingent on selling 520 loaves. Response 2 is more thorough, clearly formatted, and explicitly addresses the 450-loaf contingency with correct math, making it more useful despite slightly more length. | google/gemini-3.1-pro-preview: A — Both responses correctly calculate the profits for both scenarios. Response 1 is slightly better formatted and provides a more nuanced, business-focused recommendation. / Both models correctly calculate the profits for both scenarios and provide a clear recommendation. Model 2 is slightly better formatted and easier to read for a non-technical professional. | deepseek/deepseek-v4-pro: tie — Both responses correctly calculate profits (£60 current, £146 under deal at 520 loaves), show workings, and give a clear recommendation. No explicit constraints were broken. / Both responses accurately compute current and deal profits, showing £60 vs £146, and recommend taking the deal. No instructions broken.
tiePlanningReasoning · 8.5 v 8Both provide clear 7-day plans with owners per line and a flagged risk. Response 1 is more detailed but slightly verbose; Response 2 is more…
Plan the launch week for a small online course (already recorded). Resources: one founder, a part-time VA (10h), email list of 2,000, £300 ad budget. Produce a 7-day plan, one line per day, each line naming the owner. Flag the single riskiest dependency.
- **Day 1 — Founder:** Open enrolment with launch email, sales page check, organic social posts, and £50 test ads; **VA:** verify links/tracking and schedule support responses. - **Day 2 — Founder:** Host a short live Q&A/demo and email replay/FAQ; **VA:** collect questions, update FAQ, and monitor inbox (2h). - **Day 3 — Founder:** Send value-led case study/testimonial email and optimise ads toward best-performing audience; **VA:** compile early buyer feedback and engagement metrics (1.5h). - **Day 4 — Founder:** Publish objection-handling content (“who it’s for/not for”) and send segmented follow-up to non-openers; **VA:** schedule posts and check checkout issues (1.5h). - **Day 5 — Founder:** Run a mid-launch urgency email with bonus/reminder, shift remaining ad spend to strongest creative; **VA:** prepare FAQ/support templates and report conversions (1.5h). - **Day 6 — Founder:** Send “48 hours left” email, post personal founder story/results, and answer prospects directly; **VA:** monitor comments/inbox and resend to non-openers where appropriate (2h). - **Day 7 — Founder:** Send final-day morning and last-chance emails, make final live appearance/Q&A, close enrolment, and thank buyers; **VA:** monitor checkout/support until close and export results (1.5h). **Riskiest dependency:** The sales page, checkout, email links, and conversion tracking must work flawlessly before Day 1; with a small list and only £300 in ads, broken purchase flow or tracking would waste the limited launch window and budget.
**7-Day Launch Week Plan** - **Day 1:** Founder finalizes sales page, pricing, and email sequence; VA sets up basic tracking links. - **Day 2:** Founder writes and schedules cart-open email + social posts; VA segments list and preps ad creatives. - **Day 3:** Founder sends cart-open email to 2,000-list and goes live briefly; VA launches £100 test ads. - **Day 4:** Founder engages replies/comments and records one short urgency video; VA monitors ads and boosts winners with £100. - **Day 5:** Founder sends value + testimonial email; VA handles customer questions and updates FAQ. - **Day 6:** Founder sends penultimate reminder + limited bonus; VA allocates final £100 ad spend to top performer. - **Day 7:** Founder sends cart-close email and closes checkout; VA compiles basic sales/metrics summary. **Single riskiest dependency:** Email list deliverability + open rates (if the 2,000-list underperforms, ads and VA time cannot fully compensate).
anthropic/claude-sonnet-5: tie (flipped on swap) — Both provide clear 7-day plans with owners per line and a flagged risk. Response 1 is more detailed but slightly verbose; Response 2 is more concise and structured, with a plausible risk flag. Both meet constraints equally well. / Both meet format with one line per day and named owners, flag a risk. Response 1 is more concise and clearer; Response 2 has denser lines and slightly exceeds VA's 10h budget when summed, a minor inconsistency. | google/gemini-3.1-pro-preview: tie (flipped on swap) — Response 1 is superior because it explicitly tracks the VA's 10 hours to ensure feasibility, and its launch strategy is more realistic (opening the cart on Day 1 rather than Day 3). Both followed all formatting constraints perfectly. / Both responses follow all instructions perfectly, including the constraints of one line per day and naming the owners. Both provide practical, realistic plans that account for the limited resources and budget. Response 1 is slightly more concise, while Response 2 tracks VA hours explicitly. | deepseek/deepseek-v4-pro: tie (flipped on swap) — Response 1 details realistic, timed tasks using exact ad spend and VA hours (total 10h) and correctly flags pre-launch tech readiness as risk. Response 2 has founder finalizing sales page on Day 1, which is unrealistic; it’s less granular and overlooks critical checks. / Both plans meet all constraints: 7 days, one line per day, owner named, risk flagged. Response 1 clearly allocates the £300 ad budget; Response 2 precisely tracks the 10 VA hours. Each is correct, useful, and clear. No factual errors. They are equally strong.
tieTrade-off analysisReasoning · 9.2 v 8.8Both responses are accurate, cover all required elements within word limit, and offer similar quality recommendations with clear structure.…
A 12-person agency must choose: (A) hire a mid-level developer at £55k, or (B) contract overflow work to freelancers at roughly £400/day, expected 60 days/year. Give a recommendation in under 200 words covering cost, flexibility, quality risk, and one non-obvious factor.
**Recommendation: choose freelancers for now.** **Cost:** A £55k salary typically becomes ~£65k–£75k once employer NI, pension, equipment, software, recruitment, and management overhead are included. Freelancers at £400 × 60 days cost about **£24k/year**—materially cheaper unless demand rises above roughly 160–180 days annually. **Flexibility:** Contracting matches spend to workload and avoids carrying fixed cost during quieter periods. It also allows access to specialist skills rather than committing to one generalist. **Quality risk:** Freelancers can create inconsistency, weaker product context, and dependency on external availability. Mitigate with a vetted bench, clear acceptance criteria, code review, documentation, and ensuring an internal owner retains architectural responsibility. **Non-obvious factor:** In a 12-person agency, the bottleneck may be *senior review and client coordination*, not coding capacity. Freelancers can increase review/project-management load and erode the apparent savings. Track whether current staff have enough capacity to brief, review, and integrate external work. Hire only if overflow is sustained, predictable, and requires deep client/product knowledge for most of the year.
**Recommendation: Choose (B) freelancers.** **Cost:** Freelancers total ~£24k/year (400 × 60). A permanent hire costs £55k salary plus ~25-40% on-costs (NI, pension, benefits, equipment, recruitment), easily £70k+. Freelancers win on pure spend for this volume. **Flexibility:** Freelancers scale to actual overflow and stop when work drops; a hire is a fixed annual commitment regardless of demand. **Quality risk:** A vetted mid-level employee offers more consistency, context retention, and accountability. Freelancers vary more and may need tighter briefs/QA, raising short-term management overhead. **Non-obvious factor:** Opportunity cost of management attention—a full-time hire absorbs ongoing 1:1s, career development, and integration into a 12-person team, while freelancers stay transactional and free senior time for billable or growth work. For pure overflow at ~60 days, freelancers are the lower-risk, lower-cost default unless you foresee sustained full-time need or critical domain knowledge that must stay in-house. Re-evaluate if volume consistently exceeds ~100+ days.
anthropic/claude-sonnet-5: tie — Both responses are accurate, cover all required elements within word limit, and offer similar quality recommendations with clear structure. Response 2's non-obvious factor is slightly less novel but still valid; both are equally useful and concise. / Both cover cost, flexibility, quality risk, and a non-obvious factor within limits, with sound figures and clear recommendations. Response 2 slightly more nuanced on breakeven point and bottleneck insight, but both are comparable in quality and concision. | google/gemini-3.1-pro-preview: A — Both responses follow all instructions perfectly and are under the 200-word limit. Response 1 provides a slightly more insightful non-obvious factor regarding the bottleneck of senior review in a small agency, making it marginally more useful. / Both responses perfectly follow instructions and constraints. Response 2 wins slightly because its non-obvious factor (senior review bottleneck) is exceptionally insightful for a small agency context, making it slightly more useful. | deepseek/deepseek-v4-pro: tie — Both accurately cover cost, flexibility, quality risk, and a non-obvious factor; clear, concise, useful. No constraint breached. / Both responses are accurate, cover all required aspects, are under 200 words, and are clear and useful. They offer similar recommendations with slightly different non-obvious factors, both valid. No clear winner.
tieMaths word problemReasoning · 9.7 v 9.7Both responses are correct, follow identical logic, and arrive at 10:25 with clear working. Response 2 has a minor mislabeling (40×15 instead of…
A tank holds 2,400 litres. Pump A fills at 40 L/min. Pump B drains at 25 L/min. A runs from 09:00. B accidentally switches on at 09:20. At 10:00 B is switched off. At what time is the tank full? Show your working.
Assuming the tank is empty at 09:00: - From 09:00 to 09:20, Pump A fills alone for 20 minutes: \[ 40 \times 20 = 800\text{ L} \] - From 09:20 to 10:00, both pumps run for 40 minutes. Net filling rate: \[ 40 - 25 = 15\text{ L/min} \] Amount added: \[ 15 \times 40 = 600\text{ L} \] - Amount in tank at 10:00: \[ 800 + 600 = 1400\text{ L} \] - Remaining capacity: \[ 2400 - 1400 = 1000\text{ L} \] After 10:00, only Pump A runs at 40 L/min. Time to fill remaining amount: \[ \frac{1000}{40} = 25\text{ minutes} \] Therefore, the tank is full at **10:25**.
From 09:00 to 09:20, only pump A operates: \[ 20 \times 40 = 800 \text{ L} \] Tank contains \(800\) L at 09:20. From 09:20 to 10:00, both pumps operate (net rate \(40-25=15\) L/min): \[ 40 \times 15 = 600 \text{ L} \] Tank contains \(800+600=1400\) L at 10:00. Remaining volume: \[ 2400-1400=1000 \text{ L} \] After 10:00 only pump A operates, so time required is \[ 1000/40=25 \text{ min}. \] Thus the tank is full at 10:25. Final answer: \(10:25\)
anthropic/claude-sonnet-5: tie — Both responses are correct, follow identical logic, and arrive at 10:25 with clear working. Response 2 has a minor mislabeling (40×15 instead of 15×40) but result correct. Both equally clear and useful. / Both responses correctly compute identical steps and arrive at 10:25, with clear working. Response 2 explicitly states the empty-tank assumption, slightly improving clarity, but both are essentially equivalent in quality. | google/gemini-3.1-pro-preview: tie — Both responses correctly calculate the time the tank is full, showing clear, step-by-step working. They are equally accurate, concise, and easy to follow. / Both responses correctly solve the math problem, show clear step-by-step working, and arrive at the correct final time of 10:25. Response 2 explicitly states the assumption that the tank starts empty, but both are excellent. | deepseek/deepseek-v4-pro: tie — Both responses are factually correct, show clear step-by-step working, and meet all instructions. They are equally useful and concise. / Both responses correctly calculate 10:25, show clear step-by-step working, are useful to a non-technical professional, and fully meet all task requirements. Neither violates any explicit constraint.
tieCatch the flawReasoning · 9.3 v 8.8Both correctly identify correlation/causation issue and non-linear extrapolation flaw, within word limit, clear and concise. Response 2 adds…
Find the flaw in this argument and explain it in under 100 words: 'Our churn dropped from 6% to 3% after we introduced the loyalty programme in March, so the programme cut churn in half. We should double the loyalty budget to cut churn to 1.5%.'
The argument confuses correlation with causation: churn may have fallen because of seasonality, pricing changes, customer mix, or other simultaneous factors—not necessarily the loyalty programme. It also assumes the effect is linear: even if the programme caused some reduction, doubling spend may not halve churn again because benefits can diminish or the remaining churn may be harder to prevent. A controlled test or credible comparison group is needed.
**Correlation ≠ causation; non-linearity.** The drop after March doesn’t prove the programme caused it—other factors (seasonality, competitors, pricing, etc.) may explain the change. Even if it did, halving churn once doesn’t mean doubling spend halves it again; effects typically show diminishing returns, not linear scaling. The argument commits post-hoc and false-extrapolation fallacies.
anthropic/claude-sonnet-5: tie — Both correctly identify correlation/causation issue and non-linear extrapolation flaw, within word limit, clear and concise. Response 2 adds fallacy names, slightly more precise; Response 1 slightly more thorough on confounders. Essentially equal quality. / Both correctly identify correlation/causation and linear extrapolation flaws, within word limit, clear and concise. Response 2 adds a slightly more actionable suggestion (controlled test), Response 1 is more concise with bolded summary. Both are equally strong overall. | google/gemini-3.1-pro-preview: A — Both responses correctly identify the two main flaws (correlation vs. causation and linear extrapolation) and stay well under the 100-word limit. Response 1 is slightly more cohesive and professional in its tone, making it better suited for a business context. / Both responses correctly identify the flaws and stay under the word limit. Response 2 is better suited for a non-technical professional as it avoids academic jargon (like 'post-hoc fallacies') and provides a practical recommendation. | deepseek/deepseek-v4-pro: tie (flipped on swap) — Response 1 clearly explains both flaws without jargon, making it more accessible to non-technical readers. Response 2 is correct but uses technical fallacies names that may hinder understanding. Both adhere to word limit. / Both correctly identify correlation/causation and non-linearity. Response 2 adds practical advice, slightly more useful for non-technical users, but both are accurate and clear. Score tied at 9.
tieMessy to JSONExtraction · 9.7 v 9.5Both correctly extract fields; Response 1 is on a single line matching 'return only valid JSON' more concisely, while Response 2 uses multi-line…
Extract to JSON with keys name, company, phone, email, intent: 'hiya - jules moreau here from Atelier Nord (the lighting people). best number is 07700 900 457, or jules@ateliernord.fr. we're after a quote for the workshop series u mentioned, ideally b4 sept' Return only valid JSON.
{"name":"Jules Moreau","company":"Atelier Nord","phone":"07700 900 457","email":"jules@ateliernord.fr","intent":"Requesting a quote for the workshop series, ideally before September."}
{ "name": "Jules Moreau", "company": "Atelier Nord", "phone": "07700 900 457", "email": "jules@ateliernord.fr", "intent": "quote for the workshop series ideally before September" }
anthropic/claude-sonnet-5: tie (flipped on swap) — Both correctly extract fields; Response 1 is on a single line matching 'return only valid JSON' more concisely, while Response 2 uses multi-line formatting. Content is equivalent and accurate. Response 1 slightly more concise/clean. / Both responses correctly extract all fields into valid JSON with equivalent content. Response 2 is slightly more concise (single line, minified) matching 'return only valid JSON' instruction marginally better, but both are essentially equivalent in correctness and usefulness. | google/gemini-3.1-pro-preview: tie — Both responses correctly extract the requested information into valid JSON with the exact keys specified. Neither includes any extra text, perfectly following the negative constraint. / Both responses correctly extract the requested information into valid JSON with the specified keys. They both follow the negative constraint to return only valid JSON. Response 1 is pretty-printed, while Response 2 is minified, but both are perfectly acceptable. | deepseek/deepseek-v4-pro: tie — Both responses correctly extract all fields, follow instructions exactly, and are clear and useful. No constraint violations. / Both correctly extract all fields into valid JSON, follow the instruction precisely, and are equally useful and concise for a non-technical professional. No constraint violations.
tieTable from proseExtraction · 9.7 v 9.2Both nearly identical and correct; Response 1 includes explicit + sign for positive change, slightly clearer formatting for sign consistency with…
Turn this into a markdown table with columns Product, Q1, Q2, Change: 'The Alpha line did 14,200 units in Q1 and 15,800 in Q2. Beta slipped from 9,100 to 7,400. The new Gamma launched mid-Q2 with 3,300 units (no Q1 sales). Delta held flat at 5,000 both quarters.' Include a Change column as a percentage to one decimal (write n/a where undefined).
| Product | Q1 | Q2 | Change | |---|---:|---:|---:| | Alpha | 14,200 | 15,800 | +11.3% | | Beta | 9,100 | 7,400 | -18.7% | | Gamma | 0 | 3,300 | n/a | | Delta | 5,000 | 5,000 | 0.0% |
| Product | Q1 | Q2 | Change | |---------|------|------|--------| | Alpha | 14,200 | 15,800 | 11.3% | | Beta | 9,100 | 7,400 | -18.7% | | Gamma | 0 | 3,300 | n/a | | Delta | 5,000 | 5,000 | 0.0% |
anthropic/claude-sonnet-5: tie (flipped on swap) — Both nearly identical and correct; Response 1 includes explicit + sign for positive change, slightly clearer formatting for sign consistency with negative values. / Both correctly compute percentages and format tables identically in content; Response 2 adds alignment and a '+' sign, a minor stylistic difference but not required or forbidden. Both fully meet task requirements. | google/gemini-3.1-pro-preview: A — Both responses follow all instructions perfectly and calculate the percentages correctly. Response 1 is slightly better because it right-aligns the numerical columns and includes a plus sign for the positive change, making it easier to read. / Both responses follow all instructions perfectly and calculate the percentages correctly. Response 2 is slightly better because it right-aligns the numerical columns and includes a '+' sign for the positive change, making it easier to read. | deepseek/deepseek-v4-pro: tie — Both tables accurately present the data, compute percentages to one decimal, and use n/a for undefined. They follow all instructions and are equally clear and useful. / Both are factually accurate, follow instructions exactly, present data clearly. No constraint breaches. Identical scores.
Grok 4.5Contact detailsExtraction · 7 v 7.7Both correctly list all four people, but Response 2 provides a more accurate role for Priya Shah (Legal sign-off) instead of marking it unknown…
List every person mentioned below with their role and organisation, one line each in the format Name - Role - Org. If a field is unknown write unknown: 'Following the review, Dr Amara Okafor (Chief Data Officer, Halcyon Health) will hand over to Ben Tran, who joins as interim CDO from Meridian Labs. Legal sign-off sits with Priya Shah at Halcyon; the external audit remains with T. Lindqvist.'
Dr Amara Okafor - Chief Data Officer - Halcyon Health Ben Tran - interim CDO - Meridian Labs Priya Shah - unknown - Halcyon Health T. Lindqvist - external audit - unknown
Dr Amara Okafor - Chief Data Officer - Halcyon Health Ben Tran - interim CDO - Meridian Labs Priya Shah - Legal sign-off - Halcyon T. Lindqvist - external audit - unknown
anthropic/claude-sonnet-5: B — Both correctly list all four people, but Response 2 provides a more accurate role for Priya Shah (Legal sign-off) instead of marking it unknown, better fulfilling the task's intent while remaining concise and correctly formatted. / Both correctly list Okafor, Tran (though Org ambiguous, both same), Lindqvist. Response 1 correctly infers Priya Shah's role as legal-related, while Response 2 marks it unknown despite explicit text stating her legal role, making Response 1 more accurate. | google/gemini-3.1-pro-preview: B — Response 2 accurately captures Priya Shah's role based on the text ('Legal sign-off') and uses the exact organization name provided ('Halcyon'). Response 1 incorrectly lists her role as unknown and infers 'Health' for her organization. / Response 1 correctly identifies Priya Shah's role as 'Legal sign-off' based on the text, whereas Response 2 incorrectly lists it as 'unknown'. | deepseek/deepseek-v4-pro: tie (flipped on swap) — Response 1 correctly uses unknown for Priya Shah's role, but both miss that Ben Tran's organisation is Halcyon Health, not Meridian Labs, and T. Lindqvist's role should be unknown. Response 2 has additional errors: Priya's role is not 'Legal sign-off' and T. Lindqvist's role is not 'external audit'. Response 1 is more accurate and useful. / Response 1 correctly lists all persons except Ben Tran's org (should be unknown, not Meridian Labs) but handles other fields accurately. Response 2 additionally misses Priya Shah's known role, reducing usefulness. Both breach the explicit 'unknown' rule for Ben Tran.
tieKey factsExtraction · 8.3 v 9.5Both extract 5 accurate bullets under 15 words each, covering key facts. Response 1 omits Dock Road detail, Response 2 omits it too but mentions…
Extract exactly 5 key facts as bullets (each under 15 words) from: 'The council approved the harbour redevelopment on 12 June by 7 votes to 2. Phase one, costing £4.6m, begins in October and includes the fish market roof. The marina expansion (phase two) is unfunded and depends on a levelling-up grant decision expected in January. Local traders get first refusal on the 14 new retail units. Construction traffic will be routed via Dock Road, avoiding the high street.'
- Council approved harbour redevelopment on 12 June, voting 7–2. - Phase one costs £4.6m and begins in October. - Phase one includes repairs to the fish market roof. - Marina expansion depends on a levelling-up grant decision in January. - Local traders receive first refusal on 14 new retail units.
- Council approved harbour redevelopment on 12 June by 7-2. - Phase one costs £4.6m and begins in October. - Phase one includes the fish market roof. - Phase two unfunded, awaits January grant decision. - Local traders get first refusal on 14 retail units.
anthropic/claude-sonnet-5: tie — Both extract 5 accurate bullets under 15 words each, covering key facts. Response 1 omits Dock Road detail, Response 2 omits it too but mentions phase two explicitly. Both nearly identical quality; minor omission of construction traffic detail in both. / Both extract 5 accurate bullets under 15 words, omitting construction traffic detail but that's acceptable given 5-fact limit. Both equally clear and correct; minor wording differences don't favor either significantly. | google/gemini-3.1-pro-preview: tie (flipped on swap) — Both responses follow all constraints perfectly. Response 2 is better because it captures the 'unfunded' detail of phase two, and Response 1 hallucinates the word 'repairs' which was not in the source text. / Both responses follow all instructions perfectly, extracting exactly 5 facts with each bullet under 15 words. They are both highly accurate, concise, and useful. | deepseek/deepseek-v4-pro: B — Response 1 incorrectly adds 'repairs' to bullet 3, not in source. Response 2 is fully accurate and concise. Both meet word and bullet constraints. / Response 2 adds 'repairs' to bullet 3, inventing detail not in the text, so it is factually inaccurate. Response 1 is fully correct and follows all constraints exactly.
tieSchema complianceExtraction · 10 v 10Both responses are correct, complete, properly formatted JSON matching the schema exactly, with only whitespace differences.
Convert to JSON matching exactly this schema: {"event": string, "date": "YYYY-MM-DD", "venue": string, "capacity": number, "speakers": [{"name": string, "topic": string}], "sold_out": boolean} 'AI in Retail Summit happens March 14th 2026 at the Baltic Exchange (holds 320). Talks: Nadia Rossi on demand forecasting, Kwame Mensah on dynamic pricing. Tickets still available.' Return only the JSON.
{"event":"AI in Retail Summit","date":"2026-03-14","venue":"Baltic Exchange","capacity":320,"speakers":[{"name":"Nadia Rossi","topic":"demand forecasting"},{"name":"Kwame Mensah","topic":"dynamic pricing"}],"sold_out":false}
{"event": "AI in Retail Summit", "date": "2026-03-14", "venue": "Baltic Exchange", "capacity": 320, "speakers": [{"name": "Nadia Rossi", "topic": "demand forecasting"}, {"name": "Kwame Mensah", "topic": "dynamic pricing"}], "sold_out": false}
anthropic/claude-sonnet-5: tie — Both responses are correct, complete, properly formatted JSON matching the schema exactly, with only whitespace differences. / Both responses are identical in content, correctly matching the schema with accurate data extraction and formatting, differing only in whitespace. | google/gemini-3.1-pro-preview: tie — Both responses perfectly follow the instructions, accurately extract the information, match the required JSON schema exactly, and return only the JSON without any extra text. / Both responses perfectly extract the information into the requested JSON schema and follow the negative constraint to return only the JSON. | deepseek/deepseek-v4-pro: tie — Both JSONs are correct, match schema, all constraints met; only formatting differs. / Both responses extract all details accurately, adhere exactly to the schema, return only JSON without extra text, and are clear and useful for a non-technical professional.
GPT-5.6 TerraFormat constraintsInstruction-following · 9.2 v 7.2Both meet sentence count, starters, and word limits. Response 1 is more accurate and detailed (mentions ATP/NADPH, Calvin cycle), while Response 2…
Describe how photosynthesis works in exactly 3 sentences. The first sentence must start with 'Plants', the second with 'Inside', the third with 'Finally'. No sentence may exceed 20 words.
Plants capture sunlight with chlorophyll, using its energy to combine water and carbon dioxide into glucose and oxygen. Inside chloroplasts, light-dependent reactions split water, producing oxygen and energy-carrying molecules called ATP and NADPH. Finally, the Calvin cycle uses ATP and NADPH to convert carbon dioxide into sugar.
Plants convert sunlight water and carbon dioxide into glucose and oxygen. Inside chloroplasts chlorophyll captures light to split water and fix carbon. Finally sugars form providing energy while oxygen releases as a byproduct.
anthropic/claude-sonnet-5: A — Both meet sentence count, starters, and word limits. Response 1 is more accurate and detailed (mentions ATP/NADPH, Calvin cycle), while Response 2 is simpler but less precise. Both clear and concise. / Both meet sentence count, starting words, and word limits. Response 2 is more scientifically detailed and accurate, while still concise and compliant. | google/gemini-3.1-pro-preview: A — Both responses follow all constraints perfectly. Response 1 is much better written, more scientifically accurate, and clearer, whereas Response 2 lacks proper punctuation (e.g., missing commas in a list) making it slightly clunky. / Both responses follow all constraints perfectly. Response 2 is slightly better written with proper punctuation and provides a more accurate, yet accessible, scientific explanation. | deepseek/deepseek-v4-pro: A — Response 1 accurately details light-dependent and light-independent stages with correct terminology, fitting constraints perfectly. Response 2 oversimplifies and contains factual imprecisions, making it less correct and useful. / Response 2 provides a more accurate, detailed, and clear step-by-step explanation of photosynthesis, mentioning light-dependent reactions and the Calvin cycle, which is more useful for a professional. Both followed all constraints.
tieWord limitInstruction-following · 9.2 v 9.7Both responses are exactly 50 words, accurate, clear, and cover the key distinctions between leases and licences equally well.
Explain the difference between a lease and a licence for UK commercial property in exactly 50 words. Count carefully - exactly 50.
A lease grants exclusive possession of premises for a term, usually creating a proprietary interest that can bind successors and is protected by statute. A licence merely permits occupation or use, without exclusive possession or an estate in land; it is generally personal, more easily terminated, and offers fewer protections.
A lease grants exclusive possession of UK commercial property for a defined term, creating a proprietary interest in land enforceable against third parties. A licence merely permits occupation or use without exclusive possession, remaining a personal contractual right that is more easily revocable and does not bind successors in title.
anthropic/claude-sonnet-5: tie — Both responses are exactly 50 words, accurate, clear, and cover the key distinctions between leases and licences equally well. / Both responses are exactly 50 words, accurate, clear, and cover the key distinctions between leases and licences equally well. Quality and correctness are essentially equivalent. | google/gemini-3.1-pro-preview: B — Both responses successfully meet the strict 50-word constraint and accurately explain the legal distinction. Response 2 is slightly better as it explicitly references 'UK commercial property' as requested in the prompt, making it more tailored to the specific context. / Both responses successfully meet the strict 50-word constraint and accurately explain the legal distinction. Response 1 is slightly better as it explicitly references 'UK commercial property' as requested in the prompt, whereas Response 2 uses the generic term 'premises'. | deepseek/deepseek-v4-pro: tie (flipped on swap) — Both responses accurately distinguish lease (proprietary interest, exclusive possession) from licence (personal permission) in exactly 50 words. No constraint violations. They are clear and useful to a professional. Scores: 10 each, tie. / Both are factually accurate and exactly 50 words. Response 1 explicitly mentions UK commercial property, better addressing the task context, while Response 2 omits UK. Both are clear and useful, but 1 is slightly more precise.
tieForbidden wordsInstruction-following · 9 v 8.3Both stay within word limits (~100 words) and avoid banned words. Both explain over-extraction and heat effects clearly. Response 1 slightly more…
Explain what makes coffee taste bitter, in 80-120 words, WITHOUT using any of these words: bitter, bean, roast, caffeine, acid. Do not use hyphenated or partial variants of them either.
Coffee can develop a harsh, drying edge when heat transforms natural sugars and proteins into dark, intense compounds. Longer exposure to high temperatures increases these substances and reduces sweetness, making the cup seem more severe. Overextraction during brewing also pulls out heavier, woody, and astringent material after the desirable flavors have dissolved. Very fine grinding, water that is too hot, or contact time that is too long can cause this. Darker preparations, stale grounds, and some processing methods may amplify the effect. Balance comes from fresher grounds, suitable water temperature, proper grind size, and stopping extraction at the right moment.
The sharp and harsh flavor in coffee arises mainly from over-extraction during the brewing process. When hot water interacts too long with the ground particles, it draws out undesirable compounds that create this taste. High water temperatures accelerate this effect. Finer grinds increase surface area, leading to more of these elements being released. The specific variety of the plant and how it was processed after harvest also contribute to the levels of such compounds. To avoid this, one should control brew time, temperature, and particle size carefully for a smoother result.
anthropic/claude-sonnet-5: tie (flipped on swap) — Both stay within word limits (~100 words) and avoid banned words. Both explain over-extraction and heat effects clearly. Response 1 slightly more detailed and useful; Response 2 slightly more concise. Roughly equal quality. / Both meet word count and avoid banned words. Response 2 offers more accurate, detailed explanation of chemical causes while response 1 is slightly vaguer. Both clear and concise, but response 2 is more informative for a non-technical reader. | google/gemini-3.1-pro-preview: A — Both responses perfectly follow the negative constraints and word count limits. Response 1 is slightly better as it addresses both the heating/preparation phase and the brewing phase, providing a more comprehensive explanation of the harsh flavors. / Both responses follow all constraints, including word count and negative constraints. Response 2 provides a more comprehensive explanation by cleverly alluding to the roasting process and overextraction, making it slightly more informative. | deepseek/deepseek-v4-pro: tie (flipped on swap) — 1 covers heat-induced chemical changes, over-extraction, grind, temperature, and staleness, clearly and accurately without forbidden words. 2 is correct but less thorough. Both follow constraints exactly. / Both responses accurately explain bitterness causes (over-extraction, heat, grind) without forbidden words, stay within 80-120 words, and are clear and useful. No constraint violations.
Grok 4.5Multi-constraintInstruction-following · 5 v 8.3R1 exceeds the 110-word limit (113 words), violating the constraint. R2 meets all constraints: word count, 3 paragraphs, single question, required…
Write a LinkedIn post about hiring your first apprentice. Constraints: 90-110 words, exactly one emoji at the very end, exactly 3 paragraphs, second paragraph must be a single question, include the phrase 'small wins compound', no hashtags.
Today, we welcomed our first apprentice to the team. It feels like a meaningful milestone—not because we have everything figured out, but because we are ready to invest in potential, curiosity, and practical learning. Building a business is often about creating opportunities before you feel fully prepared. What could change if more growing teams made room for emerging talent? We’re excited to share real work, honest feedback, and space to learn from mistakes. Our new apprentice will bring fresh perspective, while we provide support, responsibility, and a clear path to grow. small wins compound, and we believe this first step will strengthen both our team and the future talent we hope to build. 🌱
Hiring my first apprentice marked a pivotal shift in how I approach team growth and knowledge sharing. Guiding someone fresh into the field requires patience and clear communication, yet the energy they bring revitalizes daily routines and sparks fresh ideas across projects. This commitment to nurturing talent strengthens our entire operation in meaningful ways. What surprised you most when you first started mentoring an apprentice? Seeing their progress reinforces that small wins compound into lasting expertise and confidence over months of steady effort. From initial shadowing sessions to independent contributions, each milestone builds momentum that benefits everyone involved. I look forward to more such journeys ahead. 🚀
anthropic/claude-sonnet-5: B — R1 exceeds the 110-word limit (113 words), violating the constraint. R2 meets all constraints: word count, 3 paragraphs, single question, required phrase, no hashtags, one emoji at end. Both are clear and useful, but R2 fully complies. / R1 stays within 90-110 words (106) meeting all constraints. R2 exceeds word limit (113 words), violating the explicit range. Both include required phrase, question paragraph, single ending emoji, no hashtags, but R1 fully complies. | google/gemini-3.1-pro-preview: tie (flipped on swap) — Response 1 failed the word count constraint (113 words). Response 2 followed all instructions perfectly, including the word count, paragraph structure, required phrase, and emoji placement. / Response 1 violates the paragraph constraint by placing the emoji on a new line (creating a 4th paragraph). Response 2 violates the word count constraint (it is over 110 words). Both otherwise follow the formatting and content instructions well. | deepseek/deepseek-v4-pro: B — Response 2 flows better with a more personalized question, and is slightly clearer. Both are useful and follow most constraints, but 1 exceeds the word limit (113 words). / Response 1 perfectly meets all constraints, includes the phrase, and delivers high clarity and usefulness. Response 2 exceeds the 110‑word limit (112 words) and is slightly less concise, but otherwise follows instructions.
tieExact structureInstruction-following · 9.5 v 9Both meet structure and word-count constraints correctly. Response 1 has clean spacing between sections improving readability, while Response 2…
Produce a checklist for closing a small retail shop at night with EXACTLY this structure: a title line, then 3 sections headed 'Front of house', 'Back office', 'Security', each containing exactly 4 numbered items, each item 3-8 words. Nothing else before or after.
Nightly Retail Shop Closing Checklist Front of house 1. Straighten displays and restock shelves 2. Clean counters and fitting rooms 3. Empty bins and remove rubbish 4. Turn off music and lights Back office 1. Count cash and reconcile tills 2. Prepare bank deposit securely 3. File receipts and daily reports 4. Check staff clock-out records Security 1. Lock stockroom and office doors 2. Arm alarm system correctly 3. Check all windows are secured 4. Lock entrance and rear doors
Closing Checklist for Small Retail Shop Front of house 1. Lock all entrance doors securely 2. Turn off display lights completely 3. Sweep and mop main floors 4. Empty trash bins thoroughly Back office 1. Balance cash register drawer 2. Secure daily sales reports 3. Shut down computer systems 4. Lock filing cabinets tightly Security 1. Set the alarm system 2. Check all window locks 3. Verify safe is secured 4. Activate overnight cameras now
anthropic/claude-sonnet-5: A — Both meet structure and word-count constraints correctly. Response 1 has clean spacing between sections improving readability, while Response 2 lacks blank lines, making it slightly less clear. Both are accurate and useful. / Both meet structure requirements with correct sections and item counts, word counts within range. Response 2 has better formatting with blank lines improving clarity, and items feel slightly more natural/varied. Response 1's 'Empty trash bins thoroughly' is only 4 words but valid; both are compliant overall. | google/gemini-3.1-pro-preview: tie (flipped on swap) — Both responses followed all constraints perfectly. Response 1 is slightly better because its categorization of tasks is more logical (e.g., locking entrance doors is placed under Security rather than Front of house). / Both responses follow all instructions perfectly, including the exact structure, section headers, item counts, and word count constraints per item. Both provide highly relevant and useful checklists for a retail shop. | deepseek/deepseek-v4-pro: tie — Both responses perfectly match the required structure: title, three sections with exact headings, four items each, all 3-8 words, no extra text. Both are factually correct and useful. / Both responses strictly follow the required structure and word counts. Content is accurate and useful, though each has a minor typo ('Sweep' vs 'Straighten'). They are equally clear and actionable.
tieQuick answer (short)Speed & cost · 10 v 10Both responses are identical, correct, concise, one-sentence answers meeting the task requirements perfectly.
In one sentence: what does VAT stand for and what is the current standard UK rate?
VAT stands for Value Added Tax, and the current standard UK rate is 20%.
VAT stands for Value Added Tax, and the current standard UK rate is 20%.
anthropic/claude-sonnet-5: tie — Both responses are identical, correct, concise, one-sentence answers meeting the task requirements perfectly. / Both responses are identical, accurate, concise, and fulfill the one-sentence requirement correctly. | google/gemini-3.1-pro-preview: tie — Both responses are identical, perfectly accurate, concise, and follow the one-sentence constraint. / Both responses are identical, perfectly accurate, and follow the single-sentence constraint exactly. | deepseek/deepseek-v4-pro: tie — Both responses are identical, factually correct, and perfectly follow the one-sentence instruction. No constraints violated. / Both responses are identical, factually correct, and follow the one-sentence instruction exactly.
tieInstant classificationSpeed & cost · 9.2 v 9.3Both correctly classify all four items. Response 1 uses cleaner two-column format without quotes/slashes; Response 2 retains quotes and slash…
Classify each as SPAM or NOT SPAM, output only two-column lines: 'You've won a £500 voucher click here' / 'Your Tesco order 8812 is ready for collection' / 'URGENT your account will be closed verify now' / 'Minutes from Thursday's board meeting attached'.
SPAM | You've won a £500 voucher click here NOT SPAM | Your Tesco order 8812 is ready for collection SPAM | URGENT your account will be closed verify now NOT SPAM | Minutes from Thursday's board meeting attached
'You've won a £500 voucher click here' / SPAM 'Your Tesco order 8812 is ready for collection' / NOT SPAM 'URGENT your account will be closed verify now' / SPAM 'Minutes from Thursday's board meeting attached' / NOT SPAM
anthropic/claude-sonnet-5: tie (flipped on swap) — Both correctly classify all four items. Response 1 uses cleaner two-column format without quotes/slashes; Response 2 retains quotes and slash format, slightly less clean. Both follow instructions reasonably well, but Response 1 is more concise and clear. / Both correctly classify all items. Response 1 uses the original order (item/label) matching the task's implied format more closely, while Response 2 reverses column order. Both are clear and concise; Response 1 slightly better aligned with expected format. | google/gemini-3.1-pro-preview: tie (flipped on swap) — Both responses correctly classify the items and follow the formatting constraint. Response 1 is slightly cleaner by omitting the quotation marks, making it easier to read. / Both responses correctly classify the items and follow the two-column constraint. Response 1 is slightly better as it retains the original quotation marks from the prompt, making it a more exact match to the input text. | deepseek/deepseek-v4-pro: tie (flipped on swap) — Both responses correctly classify all phrases, follow the two-column line format without extra text, and are clear and concise. No constraint violations. / Both classify correctly and use two-column lines. Response 1 puts the phrase first, which is more natural for scanning; Response 2 reverses columns, slightly less intuitive for a busy professional.
tieShort translationSpeed & cost · 9.2 v 9.5Both accurate, informal translations. Response 2 uses unnecessary bold formatting and slightly awkward apostrophe contraction 't'arrives'.…
Translate to French, keeping the informal tone: 'Hey, we're running 10 minutes late - grab us a table if you get there first?'
Salut, on a 10 minutes de retard — prends-nous une table si tu arrives avant nous ?
**Salut, on a 10 minutes de retard – prends-nous une table si t’arrives avant ?**
anthropic/claude-sonnet-5: tie (flipped on swap) — Both accurate, informal translations. Response 2 uses unnecessary bold formatting and slightly awkward apostrophe contraction 't'arrives'. Response 1 is cleaner and more natural. / Both translations are accurate, informal, and natural French. Response 1 uses casual 't'arrives', response 2 slightly more formal but still informal. Both equally clear and concise; no constraints violated. | google/gemini-3.1-pro-preview: B — Both translations are accurate and maintain an informal tone. Response 2 is slightly better as 't'arrives avant' perfectly captures the casual, conversational feel of a quick text message compared to the slightly more formal 'tu arrives avant nous'. / Response 1 captures the informal tone slightly better by using the spoken contraction 't'arrives', which perfectly matches the casual nature of the original English text. | deepseek/deepseek-v4-pro: tie (flipped on swap) — Both are accurate and informal. Response 2 uses the more colloquial contraction 't’arrives' and omits the redundant 'nous', making it slightly more natural and concise. Response 1 is fully correct but a shade less streamlined. / Both translations are factually correct and keep the informal tone. Response 1 uses a contraction ('t’arrives'), Response 2 is explicit ('tu arrives'). Both meet all instructions, are useful, clear, and concise. No constraints broken.
GPT-5.6 TerraMedium generationSpeed & cost · 9.2 v 5.5Response 1 is ~160 words, close to target; Response 2 is ~200 words, exceeding the 150-word constraint noticeably. Both are accurate and…
Write a 150-word 'About us' section for a family-run garden centre in Yorkshire founded in 1987, now run by the founders' daughter, known for rare perennials and a small cafe.
Since 1987, our family-run garden centre has welcomed gardeners from across Yorkshire with friendly advice, quality plants, and a genuine love of growing. Founded as a modest nursery, we began with a few benches, hardy shrubs, and big ambitions. Today, the founders’ daughter proudly carries that spirit forward, combining traditional horticultural knowledge with fresh ideas for modern gardens. We are especially known for our carefully chosen rare perennials: unusual, dependable plants that bring colour, texture, and character to borders through every season. Whether you are an experienced plant lover or filling your first patio pot, our team is here to help you find something you will enjoy growing. After exploring the nursery, pause in our small café for homemade cake, a proper cup of tea, and a warm Yorkshire welcome. We remain rooted in family, community, and the simple pleasure of gardens for generations of local gardeners and visitors alike.
**About Us** Nestled in the heart of Yorkshire, our family-run garden centre has been a cherished local treasure since 1987. Founded with a passion for plants and community, it began as a modest dream shared by our parents, who poured their love of horticulture into every corner of the grounds. Today, that legacy continues under the care of their daughter, who brings fresh energy while honouring the values that built our reputation. We specialise in rare and unusual perennials—carefully sourced varieties that delight collectors and inspire everyday gardeners alike. From delicate woodland treasures to bold, hard-to-find blooms, our selection is a true plant-lover’s paradise. After browsing the beds and greenhouses, visitors often linger in our cosy on-site café. Here, homemade cakes, fresh coffee and light lunches are served in a relaxed setting overlooking the gardens—perfect for a well-earned pause. Whether you’re seeking something special for your borders or simply a peaceful afternoon among the plants, we look forward to welcoming you. Come and experience the warmth of a true family garden centre, rooted in Yorkshire soil for nearly four decades.
anthropic/claude-sonnet-5: tie (flipped on swap) — Response 1 is ~160 words, close to target; Response 2 is ~200 words, exceeding the 150-word constraint noticeably. Both are accurate and well-written, but Response 1 better follows the word count instruction, making it more compliant. / Both responses meet the brief accurately and clearly, mentioning founding year, daughter's leadership, rare perennials, and cafe. Both are close to 150 words, warm in tone, and useful for a non-technical reader. Quality and style are comparably strong, with no clear differentiator. | google/gemini-3.1-pro-preview: A — Response 1 followed the exact 150-word constraint perfectly and included all requested details in a warm, professional tone. Response 2 failed the word count constraint (approx 175 words). / Response 1 failed the strict 150-word constraint (it is around 180 words). Response 2 hit exactly 150 words while seamlessly incorporating all the requested details in a warm, professional tone. | deepseek/deepseek-v4-pro: A — Response 1 is exactly 150 words, perfectly concise and covers all required elements. Response 2 is longer, less concise, but still accurate and useful. Response 1 wins. / Response 1 exceeds 150 words, violating the constraint. Both are accurate and well-written, but Response 2 meets the word limit exactly while being equally useful and clear.
tieLong generationSpeed & cost · 6.7 v 6.2Both cover location, additions, avoidances, timeline, and a common mistake accurately and clearly. Response 2 uses headers for better readability…
Write a detailed 400-word beginner's guide to setting up a home compost bin: location, what to add, what to avoid, how long it takes, and one common mistake.
## Beginner’s Guide to Setting Up a Home Compost Bin Home composting turns kitchen scraps and yard waste into a dark, crumbly soil amendment that reduces trash. A outdoor bin is enough for beginners: use a purchased plastic compost bin, a wire cylinder, or a wooden enclosure. Aim for a container at least 3 feet wide, 3 feet deep, and 3 feet tall, because a modest pile holds heat and decomposes more reliably. Choose a level, well-drained spot with easy access from the kitchen. Partial shade is ideal: full sun can dry a pile quickly, while deep shade may keep it too wet. Place the bin directly on bare soil rather than concrete so worms and microbes can enter, and so excess moisture can drain. Keep it a few feet from fences, structures, and neighbors’ windows, and ensure a hose is nearby. Successful compost needs a balance of “greens,” which supply nitrogen, and “browns,” which supply carbon. Add two to three buckets of browns for every bucket of greens. Browns include dry leaves, shredded paper or cardboard, straw, small twigs, and sawdust. Greens include fruit and vegetable peels, coffee grounds, tea leaves, crushed eggshells, grass clippings, and trimmings. Chop scraps, shred cardboard, and mix materials evenly. The pile should feel like a wrung-out sponge; water dry layers, and add dry browns if it becomes soggy. Do not add meat, fish, bones, dairy products, oils, greasy foods, pet waste, diseased plants, invasive weeds with seeds, or chemically treated wood. These materials can attract pests, create odors, spread pathogens, or introduce chemicals. Avoid large quantities of one item, compacted grass clippings or wet food scraps, because they restrict airflow. Do not add plastic-coated paper, glossy magazines, coal ash, or charcoal briquettes. Small amounts of fireplace wood ash are acceptable, but use them sparingly because too much can make compost overly alkaline. With regular turning, good moisture, and a balanced mix, finished compost may be ready in two to four months. A neglected or cold pile can take six months to a year. Turn the pile with a fork every one to two weeks to introduce oxygen and move undecomposed outer material inward. Compost is ready when it is dark, earthy smelling, and unrecognizable; screen out pieces and return them to the bin. A common beginner mistake is adding too many wet greens without enough browns. The pile then becomes slimy, compacted, and smelly. Fix it by mixing in shredded leaves or cardboard, turning thoroughly, and leaving the lid open.
**Setting Up a Home Compost Bin: A Beginner's Guide** Composting turns kitchen scraps and yard waste into nutrient-rich soil amendment, reducing trash and benefiting your garden. This guide covers the essentials for beginners: choosing a location, what to add and avoid, timeline, and a key pitfall. Aim for a simple bin setup that fits most backyards or balconies. **Choosing the Right Location** Select a level, well-drained spot with partial shade—too much sun dries the pile; full shade slows decomposition. Place it near a water source for easy moistening and close enough to your kitchen for convenient scrap disposal, but away from windows or seating to minimize any odors or pests. Ensure good airflow and accessibility year-round. Elevate the bin slightly or use a base with holes for drainage. A 3x3-foot enclosed bin (wood, plastic, or wire) works well for starters; open piles suit larger spaces. Avoid low spots that flood. **What to Add** Build a balanced mix of “greens” (nitrogen-rich) and “browns” (carbon-rich) in roughly equal volumes. Greens include fruit and vegetable scraps, coffee grounds, tea bags, fresh grass clippings, and plant trimmings. Browns include dry leaves, shredded cardboard, newspaper (non-glossy), straw, wood chips, and eggshells. Chop larger items into small pieces to speed breakdown. Layer them: start with browns for aeration, add greens, and top with more browns. Keep the pile as moist as a wrung-out sponge—add water if dry or dry materials if soggy. Turn or mix every 1–2 weeks with a pitchfork to introduce oxygen. **What to Avoid** Never add meat, fish, dairy, oily foods, or bones—these attract pests and create odors. Skip pet waste, diseased plants, weeds with seeds, glossy or colored paper, coal ash, and treated wood. Citrus and onions in moderation are fine for most bins, but excess can slow microbes. Avoid plastics, metals, or synthetics entirely. **How Long It Takes** With regular turning, proper moisture, and a good green-brown balance, finished compost (dark, crumbly, earthy-smelling material) typically takes 2–6 months. Hot composting (pile heating to 130–160°F) can finish in 1–3 months; cooler passive piles may need 6–12 months. Sift unfinished bits back into a new pile. Use the mature compost as mulch or soil mix. **One Common Mistake** A frequent beginner error is creating an imbalanced or neglected pile—too many greens without enough browns leads to a slimy, smelly mess, while forgetting to turn it causes anaerobic conditions and odors. Always monitor and adjust; start small and observe. With these steps, you’ll produce free fertilizer while cutting waste. Start today, stay consistent, and enjoy the results in your garden. (Word count: 412)
anthropic/claude-sonnet-5: tie (flipped on swap) — Both cover location, additions, avoidances, timeline, and a common mistake accurately and clearly. Response 2 uses headers for better readability; both are near the 400-word target, response1 slightly under likely, response2 claims 412 words. Content quality is comparable, so it's a tie. / Both are accurate, complete, well-structured guides covering all required elements. Response 1 explicitly confirms word count (412) meeting the constraint closely; Response 2 has no count but appears slightly longer and has a grammar error ('A outdoor bin'). Response 1 is clearer and more organized with headers. | google/gemini-3.1-pro-preview: tie — Both responses provide excellent, accurate, and well-structured guides covering all requested topics. However, both fail the strict length constraint of being exactly 400 words long. / Both responses provide excellent, accurate, and useful guides covering all required topics. However, both failed to meet the exact 400-word count constraint (Response 1 is 416 words, Response 2 is 422 words). They tie on quality and usefulness. | deepseek/deepseek-v4-pro: tie (flipped on swap) — Response 1 gives a more precise green-brown ratio and a single, specific common mistake with a fix, making it slightly more accurate and focused. Both are clear and useful, but Response 1 edges out. / Both guides are accurate, cover all required points, and are clear and useful. Response 1 exceeds the 400-word limit, but content quality is equivalent.
What they cost
- api price in per 1m$1OpenRouter API12 Aug 2026stale
- api price out per 1m$6OpenRouter API12 Aug 2026stale
- api price in per 1m$2OpenRouter API12 Aug 2026stale
- api price out per 1m$6OpenRouter API12 Aug 2026stale
Speed, measured
Median response time across all 30 tasks: 2.7s (GPT-5.6 Terra) vs 7.8s (Grok 4.5).
Receipts — every citation, raw outputs
The full run — prompts, both raw outputs per task, judge reasoning, latencies, token counts and per-call cost — is downloadable as JSON. The task suite is versioned (suite-2026-07) and published on the methodology page.
- judge panelanthropic/claude-sonnet-5, google/gemini-3.1-pro-preview, deepseek/deepseek-v4-proour run (raw outputs)12 Aug 2026verified
- judge swap agreement0.6our run (raw outputs)12 Aug 2026verified
- judge swap kappa0.355our run (raw outputs)12 Aug 2026verified
- median latency ms a2720our run (raw outputs)12 Aug 2026verified
- median latency ms b7766our run (raw outputs)12 Aug 2026verified
- panel swap flip rate0.356our run (raw outputs)12 Aug 2026verified
- panel unanimous rate0.4our run (raw outputs)12 Aug 2026verified
- run cost a usd0.0611our run (raw outputs)12 Aug 2026verified
- run cost b usd0.1004our run (raw outputs)12 Aug 2026verified
- score a4our run (raw outputs)12 Aug 2026verified
- score b4our run (raw outputs)12 Aug 2026verified
- suite winswriting: a 0/b 2/tie 3 · coding: a 1/b 0/tie 4 · reasoning: a 1/b 0/tie 4 · extraction: a 0/b 1/tie 4 · instruction: a 1/b 1/tie 3 · speed cost: a 1/b 0/tie 4our run (raw outputs)12 Aug 2026verified
- tasks total30our run (raw outputs)12 Aug 2026verified
- ties22our run (raw outputs)12 Aug 2026verified
level across 30 tasks — genuinely nothing between them.
“Regression: encrypted MultiAgentV2 messages remove readable task audit trail”on GPT-5.6 Terra
“Grok 4.5 is a bigger deal than Fable 5 - YouTube”on Grok 4.5