Battles / Writing

Gemini wins 4–1

a solid win on the tasks that separated them (5 of 18 tasks were decisive) — close enough that the loser is still worth a look.

Claude vs Gemini · Writingmodels in this battle: Claude Sonnet 5 (runs Claude) vs Gemini 3.1 Pro Preview (runs Gemini)

tested 12 Aug 202618 tasks3 judges from different labs
Show the full workings
last verified 12 Aug 2026suite suite-2026-08-skills/writingjudges: openai/gpt-5.6-terra, x-ai/grok-4.5, deepseek/deepseek-v4-pro (none of them a contestant)judge protocol judge-2026-08 — constraint cap enforced in code, applied to 26 judge scores in this runnot statistically decisive — Wilcoxon signed-rank on score margins p=0.0786 (n=14); sign test on win counts p=0.375 — held to our confidence gatejudge swap-consistency (Cohen’s κ) = 0.483moderate agreement κ measures how often a judge reached the same verdict when the two answers were shown in the opposite order, corrected for agreement you would get by chance. A battle below 0.21 does not publish a winner at all; above it, this number is how much of the scoreline survived the swap test rather than a claim that the judges were reliable.

The evidence

Suite-by-suite
Writing
1134

blue = Claude wins · grey = ties · white = Gemini wins (18 tasks per suite)

Round-by-round — all 18 tasks
ClaudeCold emailWriting · 8.6 v 7.8Both meet the 120-word limit, offer a free first month, maintain a friendly professional tone, avoid obvious buzzwords, and include one clear call…
PROMPT

Write a cold email (maximum 120 words) from the owner of a 6-person UK accounting firm to a local brewery, offering a free first-month bookkeeping trial. Friendly but professional, one clear call to action, no buzzwords.

Claude · 5.8s · $0.0024

Subject: Free month of bookkeeping for [Brewery Name]? Hi [Name], I run [Firm Name], a small accountancy practice based here in [Town]. We work with a few local drinks producers and know how much admin brewing involves - VAT, excise duty, supplier invoices, all of it. I'd like to offer you a free first month of bookkeeping, no strings attached, so you can see if we're a good fit. We'll sort your books, flag anything that needs attention, and give you a clear picture of where things stand. Would you be open to a 15-minute call next week to talk through how it would work? Best, [Your Name] [Firm Name] [Phone number]

Gemini · 11.0s · $0.0158

Subject: Bookkeeping for [Brewery Name] 🍻 Hi [Name], I’m [Your Name], owner of a 6-person accounting firm here in [Town/City]. We’re huge fans of your brewery—especially the [Insert Beer Name]! I know running a brewery means juggling production, sales, and a mountain of paperwork. We help local businesses keep their numbers accurate and up to date so they can focus on their craft. To show how we can help, I’d love to offer you a free first month of bookkeeping. We’ll organise your receipts and reconcile your accounts for a month, completely free, so you can see if we’re a good fit. Would you be open to a quick 10-minute phone call next week to see how this works? Best regards, [Your Name]

JUDGE (blind, position-swapped)

openai/gpt-5.6-terra: tie (flipped on swap) — Both meet the 120-word limit, offer a free first month, maintain a friendly professional tone, avoid obvious buzzwords, and include one clear call to action. Response 1 is more brewery-specific; Response 2 explicitly states the six-person firm detail. Response 2’s emoji is slightly less professional, but not materially harmful. / Response 1 meets the sender profile, stays within 120 words, is friendly and professional, and has one clear CTA. Response 2 is concise and brewery-relevant, but it does not establish that the sender is the owner of a 6-person firm, a stated requirement. | x-ai/grok-4.5: A — Both under 120 words, one CTA, professional tone, no real buzzwords. R1 stronger on UK relevance (VAT/excise) and polish; R2 adds emoji and slight casualness reducing professionalism. / Both under 120 words, friendly/pro, one CTA, no buzzwords. R1 specifies 6-person exactly; R2 better UK tone (VAT/excise, accountancy) and tighter. Both useful/clear. | deepseek/deepseek-v4-pro: A — Response 1 is perfectly on target: warm, clear CTA, under word limit, no buzzwords. Response 2 is solid but slightly less concise, goes over 120 words, and the emoji lightly undermines professionalism. / Response 2 better addresses a brewery's specific pain points (VAT, excise duty) showing industry insight, making it more useful. Both follow all constraints, are clear, and have one CTA. Response 1 is slightly more generic.

tieProduct descriptionWriting · 9.3 v 9.3Both meet the 60–80 word limit and avoid prohibited wording. Response 1 is more directly tailored to gym-goers, clearly covers every product…
PROMPT

Write a 60-80 word product description for a stainless-steel insulated water bottle (750ml, keeps drinks cold 24h and hot 12h, leakproof, fits car cup holders). Audience: gym-goers. Do not use the phrase 'stay hydrated' or the word 'sleek'.

Claude · 5.3s · $0.0018

Fuel your workouts with this 750ml stainless-steel insulated bottle, built for gym life. Double-wall vacuum insulation keeps water ice-cold for 24 hours or coffee hot for 12, so you're refreshed pre- and post-training. The leakproof lid means it's safe in your gym bag, and the slim design slides right into car cup holders. Durable, easy to clean, and ready for daily reps—this bottle keeps pace with your grind, session after session.

Gemini · 10.7s · $0.0163

Crush your workouts with the ultimate 750ml stainless-steel water bottle. Built for the gym and your daily commute, it perfectly fits standard car cup holders. Double-wall insulation keeps your pre-workout icy cold for 24 hours or your morning coffee piping hot for 12 hours. Toss it in your gym bag without a second thought. The secure, leakproof lid guarantees zero spills. Fuel your fitness journey with a durable companion designed to handle your toughest training sessions.

JUDGE (blind, position-swapped)

openai/gpt-5.6-terra: tie (flipped on swap) — Both meet the 60–80 word limit and avoid prohibited wording. Response 1 is more directly tailored to gym-goers, clearly covers every product feature, and has stronger flow. Response 2 is also accurate and useful, but uses more generic promotional phrasing and adds unsupported “standard” cup-holder wording. / Both meet the 60–80 word requirement, include all specified features, target gym-goers, and avoid forbidden wording. Response 1 is fully grounded in the provided details and is highly clear. Response 2 is also strong, but adds unsupported claims such as “easy to clean” and “slim design.” | x-ai/grok-4.5: tie — Both hit 60-80 words, cover all specs accurately for gym audience, avoid banned terms, and are clear/useful. Nearly identical quality. / Both hit 60-80 words, cover all specs accurately for gym audience, avoid bans, and are clear/useful. Equal quality. | deepseek/deepseek-v4-pro: tie — Both descriptions are factually accurate, meet all constraints (word count, banned terms), and are clear, concise, and tailored to gym-goers. No violations. / Both accurately include specs, avoid banned terms, meet word count, and target gym-goers clearly. Response 2 adds 'easy to clean' and more gym-specific terms like 'grind' and 'daily reps', giving it a slight edge in usefulness, but overall both are excellent.

tieSummarise messy notesWriting · 8.5 v 9.7Both provide exactly five clear, director-ready bullets and accurately capture the delay, campaign disagreement, budget, legal status, and Tuesday…
PROMPT

Turn these meeting notes into a clean five-bullet summary a director could read in 20 seconds: 'ok so launch slipped again, marc says supplier issue?? maybe 3 wks. sarah wants to push the email campaign anyway but tom thinks we look stupid if we promo something not shippable. budget is fine. legal still havent signed the claims doc. next check in tues.' Exactly five bullets.

Claude · 4.4s · $0.0018

Here's a clean five-bullet summary: - **Launch delayed ~3 weeks** due to a supplier issue (per Marc) - **Marketing conflict:** Sarah wants to proceed with email campaign; Tom opposes promoting a product that isn't shippable - **Budget:** On track, no issues - **Legal:** Claims document still unsigned — outstanding blocker - **Next check-in:** Tuesday

Gemini · 8.5s · $0.0106

* **Launch Delay:** Pushed back approximately 3 weeks due to a supplier issue. * **Marketing Strategy:** Team is split on running the email campaign now (Sarah) vs. waiting until the product is shippable (Tom). * **Budget:** Currently on track with no issues. * **Legal Status:** The claims document is still pending final legal sign-off. * **Next Steps:** The team will reconvene for a status update on Tuesday.

JUDGE (blind, position-swapped)

openai/gpt-5.6-terra: B — Both provide exactly five clear, director-ready bullets and accurately capture the delay, campaign disagreement, budget, legal status, and Tuesday check-in. Response 1 slightly overstates legal as a “blocker,” while Response 2 more neutrally reflects the notes and avoids unnecessary introductory text. / Both accurately capture all key points in exactly five bullets and are clear for a director. Response 1 is slightly stronger because it provides only the requested five-bullet summary. Response 2 adds an unnecessary introductory sentence and characterizes legal as a “blocker,” which is plausible but slightly more interpretive than the notes. | x-ai/grok-4.5: tie (flipped on swap) — Both deliver accurate, concise 5-bullet summaries matching notes exactly; equally clear and useful for a director. / Both accurate and clear five-bullet summaries. R1 is pure bullets meeting the exact constraint; R2 adds preamble text. | deepseek/deepseek-v4-pro: tie (flipped on swap) — Both accurately summarize the notes in exactly five bullets, are clear and concise, and meet all constraints. No violations. / Both accurate and have five bullets. Response 1 is cleaner, presenting directly without an extra introductory line, which is more suitable for a director's quick scan.

GeminiTone rewriteWriting · 8.7 v 9.5Both are warm, responsible, accurate, and under 90 words. Response 2 most directly preserves the original facts while clearly taking…
PROMPT

Rewrite this so it is warm, takes responsibility, and keeps every fact identical, in under 90 words: 'Your order was delayed because of a system error on our side. The refund has been processed and will arrive in 3-5 days. We cannot expedite it.'

Claude · 4.6s · $0.0015

I'm really sorry for the delay with your order—it was caused by a system error on our end, and that's on us. I've gone ahead and processed your refund, which will arrive within 3–5 days. I wish I could speed that timeline up for you, but unfortunately I'm not able to expedite it. Thank you so much for your patience, and please don't hesitate to reach out if you have any other questions in the meantime.

Gemini · 9.1s · $0.0120

I am truly sorry your order was delayed. This was caused by a system error on our end, and we take full responsibility for the frustration this caused you. Your refund has already been processed and will arrive in 3-5 days. While I wish we could get your money back to you faster, we unfortunately cannot expedite it. Thank you so much for your patience and understanding while we make this right.

JUDGE (blind, position-swapped)

openai/gpt-5.6-terra: B — Both are warm, responsible, accurate, and under 90 words. Response 2 most directly preserves the original facts while clearly taking responsibility. Response 1 is also strong, but “within 3–5 days” is slightly less exact than “in 3–5 days,” and it adds an invitation to contact support. / Both are warm, accountable, clear, and under 90 words. Response 1 preserves all core facts precisely and avoids implying any change in refund timing or handling. Response 2 is also strong, but “I’ve gone ahead and processed” slightly reframes the already-processed refund as a newly completed personal action. | x-ai/grok-4.5: tie (flipped on swap) — Both under 90 words, warm, responsible. R1 slightly alters refund fact to 'I've processed'; R2 keeps all facts identical and is clearer. / Both warm, take responsibility, preserve all facts exactly, and stay under 90 words with clear concise language useful to non-experts. | deepseek/deepseek-v4-pro: B — Both responses accurately preserve all original facts, convey warmth and responsibility, and stay under 90 words. Response 2 is more concise, better structured, and uses straightforward language, which benefits a busy non-technical professional. / Response 1 accurately kept all facts identical, taking responsibility warmly. Response 2 added personal action not in original, slightly altering fact. Both under 90 words, clear and warm.

tieHeadline setWriting · 8 v 8.5Both provide exactly five numbered headlines, all under 60 characters, with no colons or questions. Response 2 is more consistently focused on…
PROMPT

Write five headlines for a blog post about small UK retailers switching from Shopify to WooCommerce. Each must be under 60 characters. Number them. No colons and no questions.

Claude · 13.4s · $0.0160

1. Why UK Retailers Are Ditching Shopify for WooCommerce 2. Small UK Shops Make the Switch to WooCommerce 3. UK Retailers Choose WooCommerce Over Shopify 4. The Great Shopify Exodus Among UK Small Businesses 5. WooCommerce Wins Over UK's Independent Retailers

Gemini · 15.4s · $0.0234

1. Why Small UK Stores Ditch Shopify For WooCommerce 2. Small UK Retailers Swap Shopify For WooCommerce 3. UK Indies Trade Shopify For WooCommerce 4. Small UK Brands Move From Shopify To WooCommerce 5. UK Small Shops Leave Shopify For WooCommerce

JUDGE (blind, position-swapped)

openai/gpt-5.6-terra: B — Both provide exactly five numbered headlines, all under 60 characters, with no colons or questions. Response 2 is more consistently focused on small UK retailers and the Shopify-to-WooCommerce switch, while Response 1 includes slightly broader phrasing such as independent retailers and small businesses. / Response 1 directly addresses small UK retailers switching from Shopify to WooCommerce in every headline and follows all format constraints. Response 2 meets the explicit format rules, but several headlines are less specific about the switch from Shopify, making it slightly less aligned and useful. | x-ai/grok-4.5: tie (flipped on swap) — Both use Why-headlines (questions) violating ban; else compliant on count/format/length. R2 punchier, more varied, useful. / Both meet all constraints with accurate, numbered headlines under 60 chars and no colons/questions. Resp1 more consistently on-topic and concise for busy pros; Resp2 slightly less focused. | deepseek/deepseek-v4-pro: tie (flipped on swap) — Both meet constraints. Response 1 uses varied, catchy phrasing; Response 2 is repetitive with similar structure. Response 1 is slightly more engaging and useful. / Both follow all constraints exactly. Response 2 has slightly more varied and engaging headlines, but Response 1 is slightly more consistently concise. Overall equal quality.

GeminiCut by halfWriting · 4.7 v 8.8Both correctly state the original as 84 words; Response 1’s 31-word count is correct but it omits consultation with partners across the region and…
PROMPT

Cut this to half its length without losing any factual content: 'We are pleased to be able to announce that, following a period of extensive consultation with our valued customers and partners across the region, we have taken the decision to extend our opening hours at the Brighton branch. From Monday 6th October, the branch will be open from 8am until 8pm on weekdays, and from 9am until 5pm on Saturdays. We very much hope that these extended hours will make it easier for our customers to visit us at a time that suits them.' State the original and new word counts.

Claude · 11.6s · $0.0114

**Shortened version:** "From Monday 6th October, our Brighton branch will extend its opening hours: 8am–8pm on weekdays and 9am–5pm on Saturdays. This follows customer consultation and aims to make visiting us more convenient." **Word counts:** - Original: 84 words - New: 31 words

Gemini · 27.0s · $0.0487

**Shortened Text:** Following extensive consultation with regional customers and partners, we are extending the Brighton branch opening hours to make visiting easier for customers. From Monday 6th October, the branch will be open weekdays from 8am to 8pm, and Saturdays from 9am to 5pm. **Word Counts:** * **Original word count:** 84 words * **New word count:** 42 words

JUDGE (blind, position-swapped)

openai/gpt-5.6-terra: B — Both correctly state the original as 84 words; Response 1’s 31-word count is correct but it omits consultation with partners across the region and “extensive” consultation, losing factual content. Response 2 is exactly 42 words, retains the consultation, regional customers and partners, dates, hours, and purpose, and is clear. / Original text is 82 words, not 84. Response 1 is 42 words rather than half (41) and softens the stated hope into a purpose, but retains most details. Response 2 is only 31 words and omits partners, regional scope, and consultation extent; it also changes the stated hope. | x-ai/grok-4.5: B — R2 halves exactly to 42 words while retaining all key facts (consultation with customers/partners, hours, purpose). R1 is shorter but drops partners/region facts and is less precise. / Resp1 exactly halves to 42 words retaining all key facts (consultation, partners/region, hours, ease). Resp2 under-cuts to 31 and drops partners/region details. | deepseek/deepseek-v4-pro: B — Response 2 retains all factual details precisely (e.g., 'extensive consultation with regional customers and partners') and achieves exactly half length. Response 1 is slightly less detailed but still accurate and clear. / Response 1 cuts to exactly 42 words (half), retains all facts, and states counts correctly. Response 2 cuts to 31 words, not half, violating the constraint. Both are accurate and clear, but 1 follows instructions precisely.

tieBad news emailWriting · 7.2 v 9.3Both accurately own the underestimation, avoid blame, and offer a concrete staffing mitigation. Response 1 is clear but totals 132 words including…
PROMPT

Write an email telling a client their project will be two weeks late because we underestimated the integration work. Own it, no blame-shifting, offer one concrete mitigation, keep it under 130 words. Do not use the word 'unfortunately'.

Claude · 6.4s · $0.0032

**Subject: Project Timeline Update – Revised Delivery Date** Hi [Client Name], I want to give you an early, honest update: the project will now be delivered two weeks later than planned, with a new target date of [new date]. This is on us — we underestimated the complexity of the integration work, and that miscalculation is entirely our responsibility. To keep things moving and minimize impact on your end, we're assigning an additional senior engineer to the integration workstream starting this week, which will let us parallelize the remaining tasks rather than run them sequentially. I'll send a short progress update every Friday until delivery so you have full visibility. I'm sorry for the disruption this causes on your side, and I appreciate your patience. Happy to discuss further if useful. Best, [Your Name]

Gemini · 12.0s · $0.0186

Subject: Update on [Project Name] timeline Hi [Client Name], I am writing to share that we will deliver your project two weeks later than originally planned, moving our target date to [Date]. We underestimated the complexity of the integration work. This was our oversight, and I take full responsibility for the delay. To ensure we hit this new deadline without compromising quality, I have reassigned two of our senior developers to focus exclusively on completing this integration starting today. We are committed to getting this right. I will send you a progress update on Friday, but please let me know if you would like to schedule a brief call to discuss this before then. Best regards, [Your Name]

JUDGE (blind, position-swapped)

openai/gpt-5.6-terra: B — Both accurately own the underestimation, avoid blame, and offer a concrete staffing mitigation. Response 1 is clear but totals 132 words including subject and sign-off, exceeding the under-130 requirement. Response 2 is within the limit, direct, professional, and gives a specific mitigation and update plan. / Both clearly own the underestimate, state the two-week delay, avoid blame and the forbidden word, and provide a concrete staffing mitigation. Both are polished and useful. Response 2 is 132 words, exceeding the 130-word limit; Response 1 is within the limit. | x-ai/grok-4.5: tie — Both own the delay fully, offer one concrete mitigation (extra senior resource), stay under 130 words, avoid 'unfortunately', and are clear/useful. Nearly identical quality. / Both own the delay fully, offer one solid mitigation, stay under 130 words, avoid banned term, and remain clear/concise for a busy pro. | deepseek/deepseek-v4-pro: tie (flipped on swap) — Both emails fully comply, own the mistake, offer concrete mitigation, are under word limit, and avoid 'unfortunately'. They are clear, concise, and useful. / Both own the error, provide concrete mitigation, and avoid 'unfortunately'. Clarity and empathy are strong. Response 2 exceeds the 130-word limit (131).

tieJob advertWriting · 8.5 v 8.6Both meet all explicit requirements: part-time bookkeeping role, 12-person UK design agency, hybrid working, 20 hours, and £32,000 pro rata, with…
PROMPT

Write a job advert for a part-time bookkeeper at a 12-person UK design agency, 20 hours a week, hybrid, £32k pro rata. Maximum 180 words. Must include the salary and the hours. No 'rockstar', 'ninja' or 'family'.

Claude · 7.7s · $0.0047

# Part-Time Bookkeeper **Location:** London (Hybrid) | **Hours:** 20 hours/week | **Salary:** £32,000 pro rata We're a 12-person design agency looking for a meticulous, reliable bookkeeper to keep our finances in order as we grow. **The role:** - Manage day-to-day bookkeeping using Xero - Process invoices, expenses, and supplier payments - Reconcile accounts and bank statements - Prepare monthly management reports for our director - Support with payroll administration - Liaise with our external accountant for VAT returns and year-end accounts - Monitor cash flow and flag any issues early **About you:** - Proven bookkeeping experience (AAT qualified or equivalent preferred) - Confident using Xero or similar cloud accounting software - Excellent attention to detail and strong organisational skills - Comfortable working independently and managing your own workload - Good communicator, able to explain financial information clearly **What we offer:** - 20 hours per week, flexible scheduling - Hybrid working — split your time between our studio and home - £32,000 pro rata salary - A friendly, collaborative small team - The chance to make the role your own Sound like a good fit? Send your CV and a short cover note explaining why this role appeals to you.

Gemini · 15.8s · $0.0252

**Part-Time Bookkeeper** We are a collaborative team of 12 at a growing UK design agency, looking for an organized, detail-oriented Bookkeeper to keep our finances running smoothly. You’ll manage our day-to-day accounts, including invoicing, reconciling bank feeds, processing expenses, and preparing VAT returns. You'll work closely with our external accountants and ensure our creative team can focus on design, knowing the numbers are in safe hands. **The Details:** * **Salary:** £32,000 per annum, pro rata. * **Hours:** 20 hours per week (we are flexible on how these are structured across the week). * **Location:** Hybrid (a mix of remote working and time in our studio). **About You:** You have solid bookkeeping experience, proficiency in Xero or QuickBooks, and a strong grasp of UK accounting standards. You are proactive, approachable, and highly accurate. Ready to balance our books? Send your CV and a brief cover letter to jobs@designagency.co.uk.

JUDGE (blind, position-swapped)

openai/gpt-5.6-terra: tie (flipped on swap) — Both meet all explicit requirements: part-time bookkeeping role, 12-person UK design agency, hybrid working, 20 hours, and £32,000 pro rata, with none of the banned terms. Both are under 180 words. Response 1 is more detailed; Response 2 is more concise and equally clear for applicants. / Both meet the word limit, include salary, hours and hybrid working, and avoid banned terms. Response 2 is more complete and scannable for candidates, with clearer responsibilities, requirements and benefits. Response 1 is also strong but offers slightly less role detail. | x-ai/grok-4.5: tie (flipped on swap) — Both meet all constraints (under 180 words, include salary/hours, no banned terms) and are accurate. Resp1 is more complete with clear duties/requirements; Resp2 is concise but thinner on role details. / Both meet all constraints (under 180 words, include salary/hours, avoid banned terms), are accurate, clear, and useful for the role. | deepseek/deepseek-v4-pro: tie (flipped on swap) — Both ads include salary and hours, no forbidden words, and are under 180 words. Response 1 is better structured with clear sections, more detail on responsibilities and requirements, making it easier to scan and more useful for a busy reader. / Response 2 exceeds the 180-word limit (181 words), violating a hard constraint. Response 1 is accurate, concise, and fully compliant, making it the winner.

tieStructured explainerWriting · 9 v 8.8Both meet the three-subheading and 220-word requirements and are clear for non-technical owners. Response 1 more accurately distinguishes common…
PROMPT

Explain to a non-technical small business owner what a payment chargeback is, why it happens, and what they should do when they get one. Use exactly three subheadings. Maximum 220 words total.

Claude · 7.8s · $0.0049

# What Is a Chargeback? A chargeback happens when a customer disputes a charge on their credit or debit card, and the bank reverses the payment—pulling the money back out of your account. Unlike a regular refund, this is initiated by the customer's bank, not by you, and often comes with an extra fee. # Why It Happens Chargebacks usually occur for one of these reasons: - **Fraud** – The customer says they didn't authorize the purchase (sometimes true, sometimes a stolen card). - **Dissatisfaction** – The customer didn't get what they expected, received a damaged item, or never received it at all. - **Billing confusion** – The charge on their statement didn't look familiar, so they disputed it instead of contacting you first. - **Buyer's remorse** – Occasionally, people dispute a charge simply to get a refund without returning the product. # What to Do When You Get One 1. **Don't panic**—this is common, especially for online sellers. 2. **Gather evidence**: receipts, shipping confirmation, signed agreements, communication with the customer. 3. **Respond by the deadline** your payment processor gives you (usually 7–14 days). 4. **Submit your evidence** clearly and promptly to dispute it, if you believe the charge was valid. 5. **Learn and adjust**—if chargebacks keep happening, tighten your return policy, improve communication, or add fraud checks at checkout.

Gemini · 10.0s · $0.0137

### What is a Chargeback? A chargeback is a forced refund initiated by a customer's credit card company. Instead of asking your business for a refund, the customer disputes a charge directly with their bank. The bank immediately pulls the money from your account, along with a penalty fee, while they investigate the claim. ### Why Do Chargebacks Happen? Chargebacks usually happen for three main reasons: 1. **Fraud:** A criminal used a stolen credit card to buy something from your business. 2. **Dissatisfaction:** The customer feels the item was defective, not as described, or never arrived. 3. **"Friendly Fraud":** The customer simply forgot about the purchase, a family member made it, or they didn't recognize your business's name on their monthly bank statement. ### What Should You Do? First, don't panic. If the claim is valid (like true fraud), accept it and move on. However, if you know the charge is legitimate, you can fight it by submitting evidence to your payment processor. Provide receipts, shipping tracking numbers, or customer emails proving they received exactly what they paid for. Act quickly, as banks have strict deadlines for responses. To prevent future chargebacks, ensure your business name is clear on billing statements, use tracked shipping, and make it easy for customers to contact you for regular refunds.

JUDGE (blind, position-swapped)

openai/gpt-5.6-terra: tie (flipped on swap) — Both meet the three-subheading and 220-word requirements and are clear for non-technical owners. Response 1 more accurately distinguishes common causes, including billing confusion, and gives a stronger step-by-step response plan. Response 2 is useful but somewhat overstates that funds are immediately pulled and labels accidental non-recognition as “friendly fraud.” / Both use exactly three subheadings, stay under 220 words, and clearly explain the concept, causes, and practical response steps. Response 1 has slightly overbroad claims about immediate fund removal and fees; Response 2's 7–14-day estimate may vary by processor. Both are highly useful and accessible. | x-ai/grok-4.5: tie — Both accurate, use exactly 3 subheadings, under 220 words, clear and useful for non-technical owners. Nearly identical quality. / Both accurate, use exactly 3 subheadings, under/at 220 words, clear and useful for non-technical owners. No meaningful differences. | deepseek/deepseek-v4-pro: tie — Both responses are factually accurate, meet the subheading and word count constraints, and are clear and useful for a non-technical owner. They cover the same ground with minor structural differences, making them essentially equally effective. / Both responses are accurate, use exactly three subheadings, and stay within the word limit. They clearly explain chargebacks, reasons, and actions, catering to a non-technical owner. No factual errors, making them equally useful.

GeminiVoice matchWriting · 5 v 9.5Response 1 strongly matches the voice but has only 63 words, not 70. Response 2 is exactly 70 words and follows the requested voice: short…
PROMPT

Here is a brand's voice: short sentences, dry humour, never exclamation marks, addresses the reader as 'you', British spelling. Write a 70-word homepage intro in that voice for a company that repairs vintage watches.

Claude · 6.2s · $0.0035

Your watch has seen things. Decades, drawers, the odd disaster. We fix it properly, not quickly. Every mechanism gets opened, cleaned, understood, then put back together by someone who actually cares. No guesswork. No shortcuts. No sending it off to some mystery workshop elsewhere. You get honesty about what's wrong, what it costs, and whether it's worth doing at all. Mostly, it is.

Gemini · 17.7s · $0.0317

Your grandfather's favourite watch is broken. Again. You probably dropped it. Or perhaps you just looked at it too hard. Vintage timepieces are fragile. They are also vastly superior to modern ones. We fix them. We take the tiny, rusted pieces and make them tick. It takes time. It costs money. But you already realised that. Send us your broken heirloom. We will return it working. Until you drop it.

JUDGE (blind, position-swapped)

openai/gpt-5.6-terra: B — Response 1 strongly matches the voice but has only 63 words, not 70. Response 2 is exactly 70 words and follows the requested voice: short sentences, dry humour, direct address, British spelling, and no exclamation marks. It is also more immediately engaging for a homepage. / Response 1 is exactly 70 words, uses British spelling, addresses you, has dry humour, and clearly explains the service. Response 2 has an appropriate tone and clear message but is only 63 words, failing the explicit 70-word requirement. It is also less distinctly dry and concise. | x-ai/grok-4.5: B — Resp2 hits exact 70 words, full voice (British spelling, dry humour, short sentences, you-address, no !). Resp1 is only 63 words so breaches count; otherwise solid match. / R1 hits exact 70 words, full voice (short dry humour, British spelling, you, no !). R2 is 63 words, weaker humour and longer sentences. | deepseek/deepseek-v4-pro: B — Both nail the voice: dry humour, short sentences, 'you' address, British spelling, no exclamation marks. Response 1 is 63 words, breaking the 70-word rule; Response 2 hits 70 exactly. Otherwise equally clear, useful, and on-brand. / Response 1 nails the voice, hits exactly 70 words, and follows all constraints. Response 2 is well-written but falls short at 63 words, violating the explicit word count.

tieReply to a bad reviewWriting · 9.4 v 9.4Both directly acknowledge the 50-minute wait and lack of proactive communication, avoid a generic apology, invite the guest back once, and remain…
PROMPT

Write a public reply to this 2-star review: 'Food was fine but we waited 50 minutes for mains on a Tuesday with 6 tables occupied. Nobody said anything until I asked.' Acknowledge the specific failure, do not offer a generic apology, invite them back once, under 80 words.

Claude · 4.2s · $0.0016

Thank you for the honest feedback. A 50-minute wait for mains with only 6 tables seated is on us—that's a kitchen pacing and communication failure, not a busy-night excuse, and we should have proactively updated you instead of waiting for you to ask. We're addressing both issues with our team this week. If you're willing to give us another shot, I'd like to personally make sure your next visit reflects what we should have delivered the first time.

Gemini · 9.2s · $0.0120

You are absolutely right: waiting 50 minutes for mains on a quiet Tuesday with only six tables seated is unacceptable. Worse still is that our team failed to communicate this delay to you proactively. We are addressing this specific breakdown in kitchen pacing and floor communication with our staff immediately so it does not happen again. We would appreciate the chance to welcome you back and deliver the prompt, attentive service you should have received.

JUDGE (blind, position-swapped)

openai/gpt-5.6-terra: tie (flipped on swap) — Both directly acknowledge the 50-minute wait and lack of proactive communication, avoid a generic apology, invite the guest back once, and remain under 80 words. Response 1 is more personal; Response 2 is slightly more concise. / Both meet the under-80-word limit, acknowledge the 50-minute delay and lack of proactive communication, and invite the guest back once. Response 1 is more concise and polished for a public reply. Response 2 is also strong but slightly more conversational and includes unnecessary framing about a “busy-night excuse.” | x-ai/grok-4.5: tie — Both acknowledge the exact wait/comms failure without generic apology, invite once, and stay under 80 words. Equally clear, accurate, and useful. / Both acknowledge the exact wait/comms failure without generic apology, invite back once, and stay under 80 words. Equally clear, useful, and compliant. | deepseek/deepseek-v4-pro: tie (flipped on swap) — Response 1 is more personal and directly takes ownership with 'I'd like to personally make sure,' while still being specific, concise, and under 80 words. Response 2 is also excellent but slightly less personal. / Both responses acknowledge the wait and communication failure specifically, avoid generic apologies, include one invitation, and stay under 80 words. They are equally correct, useful, and clear. No constraints broken.

tieTurn features into benefitsWriting · 9 v 8.5Both provide three benefit-focused sentences under 20 words each. Response 1 more clearly conveys real-time automation from webhooks and is more…
PROMPT

Rewrite these three features as benefits for a small e-commerce owner, one sentence each, no more than 20 words each: '256-bit encryption', 'REST API with webhooks', '99.95% uptime SLA'.

Claude · 3.8s · $0.0015

Here are the three benefits: **256-bit encryption:** Your customers' payment details stay safe, building trust that turns visitors into repeat buyers. **REST API with webhooks:** Automatically sync orders across your tools in real time, saving hours of manual data entry weekly. **99.95% uptime SLA:** Your store stays open around the clock, so you never lose a sale to downtime.

Gemini · 11.7s · $0.0164

**256-bit encryption:** Protect your customers' sensitive payment data from hackers to build essential trust and prevent costly breaches. **REST API with webhooks:** Save hours of manual work by automatically syncing your store with your favorite shipping and inventory apps. **99.95% uptime SLA:** Never miss a sale by ensuring your online store stays reliably open to shoppers around the clock.

JUDGE (blind, position-swapped)

openai/gpt-5.6-terra: A — Both provide three benefit-focused sentences under 20 words each. Response 1 more clearly conveys real-time automation from webhooks and is more directly relevant to e-commerce operations. Both overstate uptime by implying downtime can never cause lost sales. Response 2 also overstates encryption’s ability to prevent breaches. / Both provide three clear, under-20-word benefit statements and accurately translate the features into business value. Response 2 is more concise and direct. Both slightly overstate the uptime SLA by implying downtime can never cause a lost sale, but Response 2 otherwise has stronger, more specific operational wording. | x-ai/grok-4.5: tie — Both correctly rewrite features as clear benefits in one <=20-word sentence each, highly useful and concise for the audience with no differences in quality. / Both accurately convert features to owner benefits in one sentence under 20 words each. Equally clear, useful, and compliant with all constraints. | deepseek/deepseek-v4-pro: tie (flipped on swap) — Both are accurate, concise, and follow the word limit. Response 1's 'turns visitors into repeat buyers' more directly highlights a revenue benefit, making it slightly more useful for a busy owner. Otherwise, both are excellent. / Both responses accurately rewrite each feature as a benefit, each under 20 words, perfectly following instructions. They are clear, useful, and factual with no constraint violations.

tieConstrained abstractWriting · 9.7 v 9.7Both summaries contain exactly 40 words before the required word-count statement, accurately retain the key findings, and are clear for…
PROMPT

Summarise the following in exactly 40 words, no more, no fewer: 'A study of 1,240 UK small businesses found that those adopting automated invoicing reduced late payments by 23% on average within six months, but that firms with fewer than five employees saw no significant change, largely because their invoice volume was too low for the effect to register.' State the word count at the end.

Claude · 7.4s · $0.0067

A UK study of 1,240 small businesses found automated invoicing cut late payments by 23% within six months. However, businesses with fewer than five employees showed no significant improvement, primarily because their invoice volume was too low to register effects. **Word count: 40**

Gemini · 31.0s · $0.0550

A study of 1,240 UK small businesses found automated invoicing reduced late payments by 23% within six months. However, companies with under five employees experienced no significant change, mainly because their invoice volume was too low to register an impact. Word count: 40

JUDGE (blind, position-swapped)

openai/gpt-5.6-terra: tie — Both summaries contain exactly 40 words before the required word-count statement, accurately retain the key findings, and are clear for non-technical readers. Both omit “on average,” a minor loss of precision. Response 1 more closely preserves “fewer than five”; Response 2 is equally concise and clear. / Both summaries contain exactly 40 words before the stated word count, preserve the study size, UK context, 23% reduction, six-month timeframe, and lack of significant change for firms under five employees due to low invoice volume. Both are clear, concise, and suitable for non-technical readers. | x-ai/grok-4.5: tie — Both deliver accurate 40-word summaries of key findings, state the count, and meet all constraints equally well with clear, concise phrasing suitable for busy readers. / Both deliver accurate 40-word summaries matching source facts and end with the required count statement. | deepseek/deepseek-v4-pro: tie — Both summaries are factually accurate, exactly 40 words, state the word count, and are clear and useful. No violations. / Both summaries are factually correct, exactly 40 words, state word count, and are clear and concise. No violations.

tieTwo audiencesWriting · 8.3 v 8.3Both follow the required A/B structure and sentence counts. Response 1 adds more unsupported specifics, including automatic real-time rates…
PROMPT

Explain the same product update — 'we now support multi-currency invoicing' — twice. First for an existing customer in one sentence. Then for a finance director evaluating us, in three sentences. Label them A and B.

Claude · 6.1s · $0.0030

**A (Existing customer, one sentence):** You can now invoice your customers in their local currency directly from your account, with exchange rates and conversions handled automatically. **B (Finance director evaluating us, three sentences):** Our platform now supports multi-currency invoicing, allowing your team to bill international clients in their native currency while automatically applying real-time exchange rates. This reduces manual reconciliation work and minimizes the risk of costly conversion errors across your finance operations. For organizations managing cross-border revenue, this means tighter control over cash flow forecasting and simplified compliance with local invoicing standards.

Gemini · 7.3s · $0.0090

**A** You can now easily bill your international clients in their local currencies directly from your dashboard with our new multi-currency invoicing feature. **B** Our platform now fully supports multi-currency invoicing, allowing your business to seamlessly bill international clients in their preferred local currencies. This eliminates the need for manual exchange rate calculations and reduces friction in your global accounts receivable workflows. Ultimately, this capability empowers your finance team to scale international operations while maintaining accurate, consolidated financial reporting.

JUDGE (blind, position-swapped)

openai/gpt-5.6-terra: B — Both follow the required A/B structure and sentence counts. Response 1 adds more unsupported specifics, including automatic real-time rates, reconciliation reduction, forecasting, and compliance. Response 2 also makes unverified benefit claims, but is less specific and more clearly tailored to a finance director. / Both meet the required A/B structure and sentence counts. Response 1 is clear and audience-appropriate, though it adds unverified claims about dashboards, exchange-rate calculations, and consolidated reporting. Response 2 is similarly useful but makes stronger unsupported claims about automatic real-time rates, error reduction, forecasting, and local compliance. | x-ai/grok-4.5: tie (flipped on swap) — Both meet exact A/B structure, 1-sentence and 3-sentence counts, and convey the update accurately. Response 1 is slightly more precise and useful for a finance director on controls and compliance. / Both correctly deliver one-sentence A and three-sentence B explanations of the update, properly labeled, clear, and useful for the audiences. | deepseek/deepseek-v4-pro: tie (flipped on swap) — Both meet constraints and are accurate. Response 1's finance director explanation emphasizes cash flow forecasting and compliance, directly useful for that audience. Response 2 is clear but less tailored. / Both meet all constraints, are accurate, clear, and useful. Response 2 adds slight detail on automation and compliance, but both are equally strong.

tieRemove the fluffWriting · 8 v 8.7Both remove most subjective and unmeasurable language and explain omissions. Response 2 is more concise and maps the original claims to dropped…
PROMPT

Rewrite this so it contains no marketing filler and only checkable statements: 'Our revolutionary AI-powered platform leverages cutting-edge machine learning to deliver unparalleled insights that transform how forward-thinking businesses unlock growth at scale.' If a claim cannot be made checkable, drop it and say what you dropped.

Claude · 11.2s · $0.0078

**Rewritten (checkable core only):** "Our platform uses machine learning to generate business insights." *(This is the only claim that could theoretically be verified—by inspecting the software, its models, or its outputs. Everything else in the original was unverifiable and has been removed.)* **Dropped phrases and why:** - **"revolutionary"** — subjective superlative; no defined criteria for what counts as revolutionary. - **"cutting-edge"** — vague comparative; "cutting-edge" relative to what benchmark or date? - **"unparalleled insights"** — unverifiable superlative; no comparison group or metric specified. - **"transform how... businesses unlock growth"** — vague causal claim; no defined mechanism, metric, or before/after data. - **"forward-thinking businesses"** — flattering label with no objective definition. - **"at scale"** — undefined; no scale (users, revenue, data volume) specified. **To make the remaining claim actually checkable, you'd need to specify:** - Which machine learning techniques/models are used (e.g., "gradient-boosted decision trees," "transformer-based NLP model"). - What data the insights are based on. - What specific outputs or metrics the platform produces (e.g., "predicts customer churn with X% accuracy on Y dataset").

Gemini · 10.0s · $0.0116

**Rewritten statement:** "Our platform uses machine learning to analyze business data." **What was dropped and why:** * **"Revolutionary" and "cutting-edge":** Dropped because they are subjective opinions and cannot be objectively measured or verified. * **"Unparalleled insights":** Dropped because "unparalleled" is an absolute claim that cannot be proven (it is impossible to check every other insight in the world to verify yours are unmatched). "Insights" was changed to "analyze data" to reflect the actual, verifiable function of the software. * **"Transform how forward-thinking businesses":** Dropped because "forward-thinking" is a subjective opinion, and "transform" lacks a measurable, verifiable baseline. * **"Unlock growth at scale":** Dropped because it is vague marketing jargon. Without specific, verifiable metrics (e.g., "increases revenue by X%"), "growth at scale" cannot be fact-checked.

JUDGE (blind, position-swapped)

openai/gpt-5.6-terra: B — Both remove most subjective and unmeasurable language and explain omissions. Response 2 is more concise and maps the original claims to dropped wording more directly. Response 1 is useful but does not explicitly address “AI-powered” or “leverages,” and “business insights” remains somewhat undefined. Both rewritten claims are potentially verifiable. / Response 1 gives a plainly verifiable functional claim and clearly identifies each dropped marketing claim with sound reasons. Response 2 is thorough, but retains “generate business insights,” where “insights” is undefined and qualitative, so the rewritten claim is not fully checkable. Its extra implementation suggestions are useful but unnecessary. | x-ai/grok-4.5: A — Both correctly strip unverifiable claims and list drops with reasons. Response 1 is more thorough, adds verification tips, slightly clearer structure for busy users. / Both correctly drop unverifiable claims and rewrite to checkable core while listing drops. Response 2 is more thorough and useful with verification tips; both clear and follow instructions fully. | deepseek/deepseek-v4-pro: tie (flipped on swap) — Both responses accurately drop marketing filler and produce a checkable statement. They list dropped phrases with clear reasons. Response 1 adds helpful specifics on verification; Response 2 is slightly more concise. No constraints were broken. / Response 1 gives a concise, purely checkable statement and clearly lists dropped claims. Response 2 adds an extra section on making claims more checkable, which, while not incorrect, slightly reduces concision and was not requested. Both follow instructions, but Response 1 is more direct and useful for a busy professional.

tieSequence a launchWriting · 8.7 v 8.8Both meet the required three-day structure, labels, one-sentence bodies, £49 positioning, and sub-45-character subjects. Response 2 is clearer and…
PROMPT

Write the subject lines and one-sentence bodies for a three-email launch sequence for a £49 online course on Amazon PPC. Emails go out on days 1, 3 and 6. Each subject line under 45 characters. Label the day on each.

Claude · 6.7s · $0.0030

# Amazon PPC Course Launch Sequence ## DAY 1 **Subject:** Stop guessing with your Amazon ad budget **Body:** Our new Amazon PPC course is live—learn the exact keyword and bidding strategies we use to cut wasted ad spend and boost sales, all for £49. --- ## DAY 3 **Subject:** The £3,000 PPC mistake most sellers make **Body:** Sellers routinely burn thousands on broad-match keywords and bloated ACOS—this course shows you how to fix it in under 2 hours of video lessons. --- ## DAY 6 **Subject:** Last call: Amazon PPC course closes tonight **Body:** This is your final reminder that the £49 launch price on our Amazon PPC course ends at midnight—after that, the price goes up for good.

Gemini · 12.2s · $0.0167

**Day 1** **Subject:** Master Amazon PPC for £49 **Body:** Stop burning through your ad budget and learn the exact strategies to drive profitable Amazon sales today for just £49. **Day 3** **Subject:** Lower your ACoS starting today **Body:** For less than the cost of a few wasted clicks, this £49 course will show you exactly how to optimize your campaigns and permanently slash your ACoS. **Day 6** **Subject:** Last chance: Amazon PPC for £49 **Body:** Grab the course today to stop guessing with your daily ad spend and finally turn your Amazon PPC into a predictable profit machine.

JUDGE (blind, position-swapped)

openai/gpt-5.6-terra: B — Both meet the required three-day structure, labels, one-sentence bodies, £49 positioning, and sub-45-character subjects. Response 2 is clearer and more consistently benefit-led. Response 1 includes unsupported, overly specific claims (£3,000 mistake, "we use," price rises "for good") and implies a closing deadline not provided by the task. / Response 1 cleanly meets every format requirement and offers clear, benefit-led copy. Its “permanently slash” claim is somewhat overstated. Response 2 also meets constraints and is engaging, but introduces unsupported specifics (£3,000 mistake, closing tonight, permanent price rise) that may be inaccurate or misleading without task context. | x-ai/grok-4.5: A — Both meet all constraints with under-45-char subjects, one-sentence bodies, and day labels. R1 subjects more compelling and specific; R2 solid but slightly generic. / Both meet all constraints with under-45-char subjects, one-sentence bodies, and day labels. R2 edges out on stronger urgency and specificity for a busy pro. | deepseek/deepseek-v4-pro: tie (flipped on swap) — Both follow all constraints: correct days, subject lengths under 45 characters, one-sentence bodies, labels. Equally clear and useful for a professional. / Both follow all constraints perfectly. Response 2's subject lines (with a concrete £3,000 mistake and urgent 'closes tonight') are punchier for a busy professional, giving it a slight edge in usefulness and clarity.

GeminiWrite to a deadline word countWriting · 5.7 v 8Response 1 is accurate, actionable, and well written, but has 116 words, exceeding the 90–110 limit. Response 2 meets the word count, but its…
PROMPT

Write a LinkedIn post about why most A/B tests on small e-commerce sites never reach significance. Between 90 and 110 words. No hashtags. No 'thoughts?' at the end. Open with a claim, not a question.

Claude · 6.4s · $0.0025

Most A/B tests on small e-commerce sites are dead before they start. Here's the math nobody wants to face: if your baseline conversion rate is 2%, detecting a realistic 10% relative lift requires tens of thousands of visitors per variant. Most small sites don't get that in a quarter, let alone the two weeks everyone impatiently waits before calling a "winner." So what happens instead? Teams peek early, see a promising trend, and ship it. That's not a result — that's noise wearing a lab coat. The fix isn't fancier tools. It's discipline: test bigger changes with bigger expected effects, run fewer tests with more traffic each, and accept that not everything needs a split test to greenlight.

Gemini · 18.7s · $0.0320

Most A/B tests on small e-commerce sites are a complete waste of time. The math simply doesn't work in their favor. To reach statistical significance, you need a massive sample size and a high baseline conversion rate. Small stores lack both. If your site gets 10,000 visitors a month and converts at 2%, detecting a realistic 5% uplift requires running a test for over six months. By the time you reach significance, cookie deletion, seasonality, and changing user behavior have entirely polluted your data. Instead of obsessing over minor tweaks, small brands must focus on high-impact changes, user testing, and qualitative feedback to drive growth.

JUDGE (blind, position-swapped)

openai/gpt-5.6-terra: B — Response 1 is accurate, actionable, and well written, but has 116 words, exceeding the 90–110 limit. Response 2 meets the word count, but its six-month sample-size claim is materially understated for a 2% baseline and 5% relative uplift, and its claims about data being “entirely polluted” are overstated. / Response 2 is more accurate, actionable, and clearly explains common testing mistakes. It exceeds the 90–110 word limit at 118 words. Response 1 meets the length requirement and is useful, but its six-month estimate for detecting a 5% lift with 10,000 monthly visitors substantially understates the likely duration. | x-ai/grok-4.5: tie (flipped on swap) — Both hit 105 words, open with claims, omit hashtags/thoughts. Accurate stats and advice; #1 punchier and more precise on sample sizes needed. / Both factually accurate, clear, useful posts opening with claims and meeting style rules. R1 hits 105 words exactly in range; R2 hits 119 and exceeds limit. | deepseek/deepseek-v4-pro: B — Both posts are accurate, clear, and useful. Response 1 exceeds the word limit (118 words), but per instructions I did not deduct for this. Response 2 follows all constraints. Content quality is similar, so scores are tied. / Response 1 is accurate and meets all constraints (word count 105, no hashtags, no 'thoughts?'). Response 2 exceeds the 110-word limit at 118 words, breaking an explicit constraint.

tieFaithful compressionWriting · 9.7 v 8.8Response 1 preserves all facts and numbers exactly in three clear bullets. Response 2 preserves every number and uses three bullets, but omits…
PROMPT

Compress this into three bullets, preserving every number exactly: 'Q3 revenue was £412,000, up 8% year on year. Gross margin fell from 61% to 57% because of increased shipping costs. Headcount rose from 14 to 17, and we opened the Manchester office in August, which contributed £18,000 of the quarter's revenue.'

Claude · 3.7s · $0.0012

- Q3 revenue: £412,000, up 8% year on year - Gross margin fell from 61% to 57% due to increased shipping costs - Headcount rose from 14 to 17; Manchester office opened in August, contributing £18,000 of the quarter's revenue

Gemini · 13.8s · $0.0219

• **Q3 revenue:** £412,000 (up 8% YoY). • **Gross margin:** Fell from 61% to 57% due to shipping costs. • **Expansion:** Headcount rose from 14 to 17; the new Manchester office (opened August) contributed £18,000.

JUDGE (blind, position-swapped)

openai/gpt-5.6-terra: A — Response 1 preserves all facts and numbers exactly in three clear bullets. Response 2 preserves every number and uses three bullets, but omits that shipping costs were increased and does not specify that the £18,000 was quarter revenue, reducing factual completeness. / Both use exactly three bullets and preserve every number exactly. Response 2 more faithfully retains key qualifiers, including “increased” shipping costs and that £18,000 was quarter revenue. Response 1 is concise and clear but slightly loses those details. | x-ai/grok-4.5: tie — Both compress to exactly three bullets and preserve all numbers (£412,000, 8%, 61%, 57%, 14, 17, £18,000) with full fidelity and clarity. / Both correctly compress to exactly three bullets while preserving all numbers and key facts accurately, clearly, and usefully with no material differences. | deepseek/deepseek-v4-pro: tie (flipped on swap) — Response 1 fully preserves all details including 'increased shipping costs', while Response 2 omits 'increased'—slightly reducing precision. Both follow bullet and number constraints, but Response 1 is slightly more accurate and useful. / Both responses use exactly three bullets and preserve all numbers exactly. They are factually accurate, clear, and useful. Response 2 is slightly more faithful to the original wording, but Response 1 is equally effective and concise.

What they cost
Claude
Gemini
Measured cost of this exact 18-task run: $0.0782 (Claude) vs $0.3905 (Gemini) raw outputs12 Aug 2026
Speed, measured

Median response time across all 18 tasks: 6.4s (Claude) vs 12.0s (Gemini).

Receipts — every citation, raw outputs

The full run — prompts, both raw outputs per task, judge reasoning, latencies, token counts and per-call cost — is downloadable as JSON. The task suite is versioned (suite-2026-08-skills/writing) and published on the methodology page.

Our verdict — we ran the tasks
Gemini wins 4–1

a solid win on the tasks that separated them (5 of 18 tasks were decisive) — close enough that the loser is still worth a look.

Sentiment — what people post
Claude
Negative
2+ / 11− of 13 opinions · last 90d
Gemini
Negative
2+ / 16− of 18 opinions · last 90d
Reviewed by Robert Prime
25 years building and selling ecommerce businesses, 15+ exits. Runs MrPrime and trains companies on applied AI.
changelog: 12 Aug 2026 — first published from run #43 · suite suite-2026-08-skills/writing
Too close to call 14
raw outputs ↓