Claude Sonnet 5 vs Qwen3.7 Max for workflow automation
Claude Sonnet 5 wins, 89–83
Same 12 workflow automation tasks, marked blind by three rival labs. Claude Sonnet 5 took 8 tasks, Qwen3.7 Max took 1, 3 tied. Tested 18 Aug 2026.
Where they differed most
A three-bay garage in Wakefield describes its process: 'Denise goes through the diary at the start of the month and looks for anyone whose MOT is due, well, she looks at last year's card index and the DMS, and rings them, or texts if they said they prefer text, except we lose the ones who moved, and if they book she writes it in the diary and puts a sticker on the card. If they do not answer she tries again, sometimes. About a third of them go to Kwik Fit because we never got round to it.' Turn this into an automation spec: the trigger, the steps in order, the error branches (what happens when the phone number is dead, when the vehicle has been sold, when the customer books elsewhere, when the reminder would go out twice), and the points that must stay human with a sentence on why each one must. Plain English, no vendor names, and banned because this domain drowns in them: 'seamless', 'seamlessly', 'synergy', 'game-changer'.
# MOT Reminder Automation Spec ## Trigger Monthly, on a fixed date (matching how Denise currently starts the cycle at the start of the month): the system checks all vehicle records for an MOT due date falling within the next 4–6 weeks (the window Denise implicitly used by working a month ahead). ## Steps, in order 1. **Pull the list.** Cross-reference the card index data (now presumably migrated into one system) against the DMS to identify all customers with an MOT due in the target window. This replaces checking two separate sources with o…
**Automation Spec: MOT Reminder and Booking Process** **Trigger** The process starts automatically every morning when the system identifies any vehicle in the database whose MOT expires in exactly 30 days. **Steps in Order** 1. The system pulls the list of due vehicles and checks each customer's preferred contact method (text or phone call). 2. The system checks the appointment diary to ensure the customer has not already booked an MOT for this cycle. 3. For customers who prefer texts, the system sends an automated message containing a direct…
Task by task
| Task | Claude Sonnet 5 | Qwen3.7 Max |
|---|---|---|
| The MOT reminder nobody sends | 9 | 8 |
| The spec, as JSON | 9 | 8 |
| Where the human stays | 9 | 8 |
| The unhappy paths are the job | 8 | 9 |
| Do not automate this | 10 | 9 |
| It fired twice | 9 | 9 |
| The spreadsheet that runs the business | 8 | 8 |
| The alert that wakes someone up | 10 | 9 |
| How long it really takes | 9 | 8 |
| No API, no chance? | 9 | 8 |
| Rules, not guesses | 9 | 8 |
| Refuse the scraper | 8 | 8 |
Full receipts: Claude Sonnet 5, Qwen3.7 Max · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5
Questions people ask
Which is better for workflow automation: Claude Sonnet 5 or Qwen3.7 Max?
Claude Sonnet 5 — it scored 89/100 against 83/100 on our 12-task workflow automation suite, winning 8 tasks to 1 with 3 tied. Every answer was marked blind by three judges from three rival AI labs.
How was this tested?
Both models answered the identical published workflow automation tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.
More workflow automation head-to-heads: Claude Sonnet 5 vs GPT-5.6 Sol · GPT-5.6 Sol vs Qwen3.7 Max · Claude Sonnet 5 vs GPT-5.3-Codex · GPT-5.3-Codex vs Qwen3.7 Max · Claude Sonnet 5 vs GPT-5.6 Terra · GPT-5.6 Terra vs Qwen3.7 Max
Full ranking: Best AI for workflow automation · model pages: Claude Sonnet 5, Qwen3.7 Max