Claude Opus 4.8 vs DeepSeek V4 Pro for workflow automation
Claude Opus 4.8 wins, 84–81
Same 12 workflow automation tasks, marked blind by three rival labs. Claude Opus 4.8 took 6 tasks, DeepSeek V4 Pro took 1, 5 tied. Tested 18 Aug 2026.
Where they differed most
The process, as the office manager of a small Bristol removals firm describes it: 'A quote request comes in on the website form. Someone reads it, decides if it is within our area, and if it is we email a price from the ready reckoner. If it is a big job we go and look at it first. If they accept we take a 20% deposit and put it in the calendar, and if the deposit does not arrive within a week the slot goes back.' Reply with ONLY valid JSON — no prose before or after. An object with exactly these keys and no others: trigger (string), steps (an array of objects, each with keys step (number), action (string), automated (boolean), on_error (string)), human_decisions (an array of objects with keys decision (string) and why_human (string)), and open_questions (an array of strings — the things the description does not tell you and that must be answered before this could be built).
```json
{
"trigger": "A quote request is submitted via the website form",
"steps": [
{
"step": 1,
"action": "Receive quote request from website form",
"automated": true,
"on_error": "Log the failed submission and alert office manager to follow up manually"
},
{
"step": 2,
"action": "Read the request and decide whether the job is within the service area",
"automated": false,
"on_error": "If area cannot be determined, contact customer for clarification of location"
},
{…{
"trigger": "A quote request is submitted via the website form",
"steps": [
{
"step": 1,
"action": "Read the incoming quote request",
"automated": false,
"on_error": "Not defined"
},
{
"step": 2,
"action": "Assess if the job location is within the service area",
"automated": false,
"on_error": "Not defined"
},
{
"step": 3,
"action": "Assess if the job is considered big (if within area)",
"automated": false,
"on_error": "Not defined"
},
{…Task by task
| Task | Claude Opus 4.8 | DeepSeek V4 Pro |
|---|---|---|
| The MOT reminder nobody sends | 9 | 8 |
| The spec, as JSON | 5 | 8 |
| Where the human stays | 8 | 6 |
| The unhappy paths are the job | 8 | 8 |
| Do not automate this | 9 | 9 |
| It fired twice | 9 | 9 |
| The spreadsheet that runs the business | 9 | 8 |
| The alert that wakes someone up | 9 | 8 |
| How long it really takes | 9 | 8 |
| No API, no chance? | 9 | 8 |
| Rules, not guesses | 8 | 8 |
| Refuse the scraper | 9 | 9 |
Full receipts: Claude Opus 4.8, DeepSeek V4 Pro · judges gemini-3.1-pro-preview, gpt-5.6-terra, grok-4.5
Questions people ask
Which is better for workflow automation: Claude Opus 4.8 or DeepSeek V4 Pro?
Claude Opus 4.8 — it scored 84/100 against 81/100 on our 12-task workflow automation suite, winning 6 tasks to 1 with 5 tied. Every answer was marked blind by three judges from three rival AI labs.
How was this tested?
Both models answered the identical published workflow automation tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.
More workflow automation head-to-heads: Claude Opus 4.8 vs GPT-5.6 Sol · DeepSeek V4 Pro vs GPT-5.6 Sol · Claude Opus 4.8 vs Claude Sonnet 5 · Claude Sonnet 5 vs DeepSeek V4 Pro · Claude Opus 4.8 vs GPT-5.5 · DeepSeek V4 Pro vs GPT-5.5
Full ranking: Best AI for workflow automation · model pages: Claude Opus 4.8, DeepSeek V4 Pro