Gemini 3.1 Pro Preview vs GPT-5.6 Sol for making flashcards

The verdict

GPT-5.6 Sol wins, 9895

Same 12 making flashcards tasks, marked blind by three rival labs. GPT-5.6 Sol took 4 tasks, Gemini 3.1 Pro Preview took 1, 7 tied. Tested 16 Aug 2026.

Where they differed most

The task — Definition trap

Make 4 flashcards on 'supply and demand' where NONE of the questions begins with 'What is' or 'Define'. Each must test understanding through a scenario (max 25 words per question), with a one-line answer.

**Flashcard 1**
**Q:** Coffee prices suddenly double due to a bad harvest. How will this event affect the price and quantity of tea sold? 
**A:** Both the price and quantity of tea will increase because consumers will switch to it as a substitute good.

**Flashcard 2**
**Q:** A new tax makes lithium much more expensive to mine. How does this impact the market for electric vehicle batteries?
**A:** The supply of batteries will decrease, leading to higher market prices and a lower quantity sold.

**Flashcard 3**
**Q:** The government mandates tha…
1. **Q:** A drought destroys much of the wheat crop while demand stays constant. How will wheat’s equilibrium price and quantity change?  
   **A:** Price rises, while equilibrium quantity falls.

2. **Q:** Consumers suddenly prefer electric cars, but production capacity remains unchanged. How does the market respond?  
   **A:** Demand increases, raising both equilibrium price and quantity.

3. **Q:** A government sets rent below the equilibrium price. How will this affect the housing market?  
   **A:** Quantity demanded exceeds quantity supp…

Task by task

TaskGemini 3.1 Pro PreviewGPT-5.6 Sol
Cards from a passage9.510
Application, not recall1010
Cloze cards done right109
Too big for one card910
French vocab set1010
Spot the untestable card99
Interleaved practice set1010
From messy notes99
Reverse the cards910
Definition trap810
Admit the limit1010
Refuse the sabotage1010

Full receipts: Gemini 3.1 Pro Preview, GPT-5.6 Sol · judges claude-sonnet-5, gpt-5.6-terra, grok-4.5

Questions people ask

Which is better for making flashcards: Gemini 3.1 Pro Preview or GPT-5.6 Sol?

GPT-5.6 Sol — it scored 98/100 against 95/100 on our 12-task making flashcards suite, winning 4 tasks to 1 with 7 tied. Every answer was marked blind by three judges from three rival AI labs.

How was this tested?

Both models answered the identical published making flashcards tasks. Three AI judges from three different labs scored every answer blind against a fixed rubric; mechanically checkable rules (word limits, banned phrases) are enforced by the test harness in code. The raw outputs and judge verdicts are downloadable.

More making flashcards head-to-heads: GPT-5.6 Sol vs Qwen3.7 Max · DeepSeek V4 Pro vs GPT-5.6 Sol · GPT-5.3-Codex vs GPT-5.6 Sol · GPT-5.5 vs GPT-5.6 Sol · GPT-5.6 Sol vs Kimi K3 · GPT-5.6 Sol vs Grok 4.5

Full ranking: Best AI for making flashcards · model pages: Gemini 3.1 Pro Preview, GPT-5.6 Sol