DeepSeek V4 Pro vs Gemini 3.1 Pro Preview for HR and employment questions
DeepSeek V4 Pro wins, 90–87
Same 12 HR and employment questions tasks, marked blind by three rival labs. DeepSeek V4 Pro took 5 tasks, Gemini 3.1 Pro Preview took 2, 5 tied. Tested 18 Aug 2026.
Where they differed most
Guidance excerpt: 'Records of performance conversations should be factual and contemporaneous. They should record what was said, what was agreed, by when, and what support was offered, and should avoid opinion about the employee as a person, speculation about the cause, and any reference to health or personal circumstances beyond what the employee has volunteered and agreed is relevant.' What happened: on Tuesday 12 August, at 2pm, in the small meeting room, you met Kelly Marsden. You raised three specific things: two customer orders picked to the wrong address in the last fortnight (5 August and 11 August), a returns log not completed on four of ten shifts, and arriving 20 minutes late three times since 1 August. Kelly said she is fine, that the returns log is 'everyone', and that she has a lot on at home but did not want to say more. You agreed she will complete the returns log every shift, you will re-run the picking induction with her on 19 August, and you will review on 9 September. You did not issue any warning. Write ONLY the file note — max 130 words, nothing else, no heading beyond a date line, no advice to yourself.
12 August 2025 Met with Kelly Marsden at 2pm in the small meeting room. Discussed: two orders picked to wrong address on 5 and 11 August; returns log incomplete on four of ten shifts; and three instances of lateness since 1 August. Kelly stated she is fine, said the returns log issue is “everyone”, and mentioned she has a lot on at home but chose not to elaborate. Agreed: Kelly will complete the returns log every shift; picking induction to be re-run on 19 August; review meeting scheduled for 9 September. No warning issued.
Tuesday 12 August Met with Kelly Marsden at 2pm in the small meeting room. I raised three issues: two customer orders picked to the wrong address (5 and 11 August), the returns log not completed on four of ten shifts, and arriving 20 minutes late three times since 1 August. Kelly stated she is fine. She said the returns log issue is "everyone" and mentioned she has a lot on at home, but did not want to say more. We agreed Kelly will complete the returns log every shift. As support, I will re-run the picking induction with her on 19 August.…
Task by task
| Task | DeepSeek V4 Pro | Gemini 3.1 Pro Preview |
|---|---|---|
| Probation is not a free pass | 10 | 9 |
| Write only the message | 10 | 10 |
| The grievance about the manager | 8 | 7 |
| The decision was made in the car | 10 | 9 |
| Ask everyone or ask nobody | 7.5 | 8 |
| The off-the-record phone call | 8 | 8 |
| Write only the invitation letter | 10 | 10 |
| Precedent is not a reason | 8 | 8 |
| Write only the file note | 8 | 10 |
| The number you must not supply | 10 | 8 |
| The screenshotted group chat | 9 | 8 |
| Refuse the engineered redundancy | 9 | 9 |
Full receipts: DeepSeek V4 Pro, Gemini 3.1 Pro Preview · judges claude-sonnet-5, gemini-3.1-pro-preview, gpt-5.6-terra
Questions people ask
Which is better for HR and employment questions: DeepSeek V4 Pro or Gemini 3.1 Pro Preview?
DeepSeek V4 Pro — it scored 90/100 against 87/100 on our 12-task HR and employment questions suite, winning 5 tasks to 2 with 5 tied. Every answer was marked blind by three judges from three rival AI labs.
How was this tested?
Both models answered the identical published HR and employment questions tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.
More HR and employment questions head-to-heads: Claude Sonnet 5 vs DeepSeek V4 Pro · Claude Sonnet 5 vs Gemini 3.1 Pro Preview · DeepSeek V4 Pro vs GPT-5.5 · Gemini 3.1 Pro Preview vs GPT-5.5 · DeepSeek V4 Pro vs GPT-5.6 Sol · Gemini 3.1 Pro Preview vs GPT-5.6 Sol
Full ranking: Best AI for HR and employment questions · model pages: DeepSeek V4 Pro, Gemini 3.1 Pro Preview