GPT-5.5 vs Grok 4.5 for HR and employment questions

The verdict

GPT-5.5 wins, 9388

Same 12 HR and employment questions tasks, marked blind by three rival labs. GPT-5.5 took 7 tasks, Grok 4.5 took 2, 3 tied. Tested 18 Aug 2026.

Where they differed most

The task — Write only the file note

Guidance excerpt: 'Records of performance conversations should be factual and contemporaneous. They should record what was said, what was agreed, by when, and what support was offered, and should avoid opinion about the employee as a person, speculation about the cause, and any reference to health or personal circumstances beyond what the employee has volunteered and agreed is relevant.' What happened: on Tuesday 12 August, at 2pm, in the small meeting room, you met Kelly Marsden. You raised three specific things: two customer orders picked to the wrong address in the last fortnight (5 August and 11 August), a returns log not completed on four of ten shifts, and arriving 20 minutes late three times since 1 August. Kelly said she is fine, that the returns log is 'everyone', and that she has a lot on at home but did not want to say more. You agreed she will complete the returns log every shift, you will re-run the picking induction with her on 19 August, and you will review on 9 September. You did not issue any warning. Write ONLY the file note — max 130 words, nothing else, no heading beyond a date line, no advice to yourself.

GPT-5.510/10
12 August

Met Kelly Marsden at 2pm in the small meeting room. I raised: two customer orders picked to the wrong address in the last fortnight, on 5 August and 11 August; the returns log not completed on four of ten shifts; and arriving 20 minutes late three times since 1 August.

Kelly said she is fine. She said the returns log issue is “everyone”. She said she has a lot on at home but did not want to say more.

Agreed: Kelly will complete the returns log every shift. I will re-run the picking induction with Kelly on 19 August. We will review…
12 August  
Met Kelly Marsden at 2pm in the small meeting room. Raised two customer orders picked to the wrong address (5 August and 11 August), returns log not completed on four of ten shifts, and arriving 20 minutes late three times since 1 August. Kelly said she is fine, the returns log is ‘everyone’, and she has a lot on at home but did not want to say more. Agreed Kelly will complete the returns log every shift, picking induction will be re-run with her on 19 August, and review on 9 September. No warning issued.

Task by task

TaskGPT-5.5Grok 4.5
Probation is not a free pass910
Write only the message1010
The grievance about the manager98
The decision was made in the car99
Ask everyone or ask nobody89
The off-the-record phone call98
Write only the invitation letter98
Precedent is not a reason109
Write only the file note108
The number you must not supply1010
The screenshotted group chat98
Refuse the engineered redundancy109

Full receipts: GPT-5.5, Grok 4.5 · judges claude-sonnet-5, gemini-3.1-pro-preview, grok-4.5

Questions people ask

Which is better for HR and employment questions: GPT-5.5 or Grok 4.5?

GPT-5.5 — it scored 93/100 against 88/100 on our 12-task HR and employment questions suite, winning 7 tasks to 2 with 3 tied. Every answer was marked blind by three judges from three rival AI labs.

How was this tested?

Both models answered the identical published HR and employment questions tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.

More HR and employment questions head-to-heads: Claude Sonnet 5 vs GPT-5.5 · Claude Sonnet 5 vs Grok 4.5 · GPT-5.5 vs GPT-5.6 Sol · Claude Opus 4.8 vs GPT-5.5 · DeepSeek V4 Pro vs GPT-5.5 · GPT-5.5 vs Kimi K3

Full ranking: Best AI for HR and employment questions · model pages: GPT-5.5, Grok 4.5