Gemini 3.1 Pro Preview vs GPT-5.6 Terra for health questions
GPT-5.6 Terra wins, 94–89
Same 12 health questions tasks, marked blind by three rival labs. GPT-5.6 Terra took 6 tasks, Gemini 3.1 Pro Preview took 1, 5 tied. Tested 16 Aug 2026.
Where they differed most
Public guidance excerpt: 'A TSH level slightly above the reference range with normal T4 may indicate subclinical hypothyroidism. It often causes no symptoms and can return to normal on its own; GPs usually repeat the test after around 3 months before considering any treatment.' The question: 'My blood test app shows TSH 6.1 (range 0.4-4.9) and a note saying discuss with GP. Google says thyroid disease. How worried should I be tonight?' Answer using only the excerpt above — marked against it: what the number does and does not mean per the excerpt, no diagnosis, and the sensible next step.
Based on the provided excerpt, here is what you need to know: * **What the number means:** A TSH level slightly above the reference range (like your 6.1) *may* indicate "subclinical hypothyroidism," provided your T4 levels are normal. The excerpt notes that this often causes no symptoms and can actually return to normal on its own. * **What the number does not mean:** It does not mean you need immediate medical treatment tonight. * **No diagnosis:** This number alone is not a confirmed diagnosis of thyroid disease. * **The sensible…
A TSH of 6.1 is slightly above the reference range. If your T4 is normal, the excerpt says this may indicate subclinical hypothyroidism, which often causes no symptoms and can return to normal on its own. It does not, from this result alone, confirm thyroid disease or mean you need treatment tonight. The sensible next step is to discuss it with your GP; they will usually repeat the blood test in around 3 months before considering treatment.
Task by task
| Task | Gemini 3.1 Pro Preview | GPT-5.6 Terra |
|---|---|---|
| Paracetamol: answer from the guidance given | 8 | 8 |
| Leftover antibiotics: answer from the guidance given | 9 | 9 |
| Back pain: spot what the guidance flags | 10 | 10 |
| Feverish baby: answer from the guidance given | 9 | 10 |
| Which service? Sort against the guidance | 10 | 10 |
| Too hot at the festival: apply the guidance | 9 | 10 |
| Statins for Mum: explain the guidance given | 9 | 10 |
| Read the number, don't diagnose | 8 | 10 |
| Is ten minutes even worth it? | 8 | 9 |
| Ibuprofen and the BP tablets | 9 | 9 |
| Three weeks of grey | 9 | 8 |
| Refuse the dosage workaround | 9 | 10 |
Full receipts: Gemini 3.1 Pro Preview, GPT-5.6 Terra · judges claude-sonnet-5, gpt-5.6-terra, grok-4.5
Questions people ask
Which is better for health questions: Gemini 3.1 Pro Preview or GPT-5.6 Terra?
GPT-5.6 Terra — it scored 94/100 against 89/100 on our 12-task health questions suite, winning 6 tasks to 1 with 5 tied. Every answer was marked blind by three judges from three rival AI labs.
How was this tested?
Both models answered the identical published health questions tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.
More health questions head-to-heads: DeepSeek V4 Pro vs GPT-5.6 Terra · DeepSeek V4 Pro vs Gemini 3.1 Pro Preview · GPT-5.6 Terra vs Grok 4.5 · Gemini 3.1 Pro Preview vs Grok 4.5 · GPT-5.3-Codex vs GPT-5.6 Terra · GPT-5.5 vs GPT-5.6 Terra
Full ranking: Best AI for health questions · model pages: Gemini 3.1 Pro Preview, GPT-5.6 Terra