DeepSeek V4 Pro vs GPT-5.3-Codex for emotional support
Effectively a dead heat: 97–95
Same 12 emotional support tasks, marked blind by three rival labs. DeepSeek V4 Pro took 4 tasks, GPT-5.3-Codex took 2, 6 tied. Tested 16 Aug 2026.
Where they differed most
Your friend's dog Alfie died this morning after 14 years. Write the text message (max 60 words) you would actually send. Specific to Alfie, no fixing, no silver linings — banned phrases: 'rainbow bridge', 'in a better place', 'everything happens for a reason'. It should ask for nothing except what she wants to give.
I just heard about Alfie. I know he was your boy for 14 years, and I’m so sorry. I’m not going to ask anything of you. I’m here to hear Alfie stories, sit in the quiet, or whatever you want. No need to reply.
I’m so sorry about Alfie. Fourteen years is a lifetime of his little routines and his big heart, and I know your home feels impossibly quiet today. I’m holding you both in my heart. If you feel like it, tell me your favorite Alfie story anytime—no pressure to reply.
Task by task
| Task | DeepSeek V4 Pro | GPT-5.3-Codex |
|---|---|---|
| The dog died | 10 | 9 |
| Redundant at 55 | 10 | 10 |
| She said no advice | 10 | 10 |
| After the miscarriage | 10 | 9 |
| Driving test at 9am | 8 | 9 |
| First Father's Day after | 10 | 9 |
| The 3am calls | 10 | 10 |
| Days without speaking | 9 | 10 |
| Stop apologising | 10 | 10 |
| Not a diagnosis | 9 | 9 |
| Her promotion, your rejection | 10 | 10 |
| The 1am message | 10 | 9 |
Full receipts: DeepSeek V4 Pro, GPT-5.3-Codex · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra
Questions people ask
Which is better for emotional support: DeepSeek V4 Pro or GPT-5.3-Codex?
Effectively a dead heat: DeepSeek V4 Pro edged it 97/100 to 95/100 on our emotional support suite — too close to matter, so pick on price or the product you already use.
How was this tested?
Both models answered the identical published emotional support tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.
More emotional support head-to-heads: Claude Sonnet 5 vs DeepSeek V4 Pro · Claude Sonnet 5 vs GPT-5.3-Codex · DeepSeek V4 Pro vs GPT-5.5 · GPT-5.3-Codex vs GPT-5.5 · Claude Fable 5 vs DeepSeek V4 Pro · Claude Opus 4.8 vs DeepSeek V4 Pro
Full ranking: Best AI for emotional support · model pages: DeepSeek V4 Pro, GPT-5.3-Codex