Battles / Images

Probably GPT Image (in ChatGPT)

Gemini vs GPT Image (in ChatGPT) · Images

GPT Image (in ChatGPT) came out in front, but not by enough for us to call it proven on a suite this size. Treat it as the way to bet, not as a settled result.

24
of 6 decided · 4 tied

GPT Image (in ChatGPT) took 4 of the 6 tasks that had a clear winner (Gemini 2, GPT Image (in ChatGPT) 4). The judge could pick a winner on 6 of 10 tasks; on the other 4 it could not tell them apart. That is a lean, not a proven win — at this sample size we cannot rule out chance, so we are not calling it decisive.

So which should you pick?

Worth knowing: Gemini answered noticeably faster.

Gemini answered faster - 9.3s against 183.1s
tested 17 Jul 202610 tasks1 judge — predates our three-lab panel
Show the full workings
last verified 17 Jul 2026suite images-2026-07judges: anthropic/claude-sonnet-5 (none of them a contestant)judge protocol judge-2026-07 — predates enforcement of the constraint cap; the rubric stated it and nothing applied itnot statistically decisive — Wilcoxon signed-rank on score margins p=0.3717 (n=9); sign test on win counts p=0.6875 — held to our confidence gate

The evidence

Suite-by-suite
Product shots
110
Text in image
110
Illustration
002
Photorealism
011
Diagrams
011

blue = Gemini wins · grey = ties · white = GPT Image (in ChatGPT) wins (2 tasks per suite)

Round-by-round — all 10 tasks
GeminiProduct shotProduct shots · 8 v 7Pass 1 (A first): Both meet the prompt well; Image 1 has slightly better lighting gradient and composition with more natural top-left key light…
PROMPT

A studio product photograph of a matte-black insulated coffee flask on a light grey seamless background, soft top-left key light, subtle reflection under the flask, no text, no watermark, e-commerce quality.

Gemini · 7.4s · $0.0685
Gemini — Product shot
GPT Image (in ChatGPT) · 133.9s · $0.2940
GPT Image (in ChatGPT) — Product shot
JUDGE (blind, position-swapped)

Pass 1 (A first): Both meet the prompt well; Image 1 has slightly better lighting gradient and composition with more natural top-left key light, though flip-cap design resembles bottle more than flask. Image 2 is classic flask shape but flatter lighting. | Pass 2 (B first): Both fit the brief; image 2 looks more like a classic insulated flask with better lighting gradient and subtle reflection, though it resembles a water bottle with handle/straw. Image 1 is clean but shadow is harsher and lighting flatter.

tieLifestyle productProduct shots · 7 v 7Pass 1 (A first): Image 1 shoes are dirty, not clean white as requested. Image 2 shows cleaner white shoes, nicer dawn light, wet reflective…
PROMPT

A pair of white running shoes on a wet city pavement at dawn, shallow depth of field, natural morning light, photorealistic, no people, no text.

Gemini · 8.5s · $0.0685
Gemini — Lifestyle product
GPT Image (in ChatGPT) · 193.0s · $0.2933
GPT Image (in ChatGPT) — Lifestyle product
JUDGE (blind, position-swapped)

Pass 1 (A first): Image 1 shoes are dirty, not clean white as requested. Image 2 shows cleaner white shoes, nicer dawn light, wet reflective pavement, better composition and aesthetic appeal, though pavement is more concrete than asphalt. | Pass 2 (B first): Image 2 shows more convincing wet pavement texture, natural dawn lighting, and realistic photographic quality. Image 1 has smoother, slightly artificial-looking shoes and less convincing wetness. Both lack text/people as required. Image 2's Nike logo isn't legible text, acceptable.

GeminiText rendering — posterText in image · 8 v 6Pass 1 (A first): Image 1 has cleaner Swiss minimalism with flat pear icon and correct text hierarchy including em-dashes; Image 2's textured pear…
PROMPT

A minimalist concert poster with the exact text "MIDNIGHT ORCHARD — 14 AUGUST — ROUNDHOUSE" in a clean sans-serif, dark navy background, a single stylised pear illustration, Swiss design style.

Gemini · 13.8s · $0.0672
Gemini — Text rendering — poster
GPT Image (in ChatGPT) · 154.9s · $0.2930
GPT Image (in ChatGPT) — Text rendering — poster
JUDGE (blind, position-swapped)

Pass 1 (A first): Image 1 has cleaner Swiss minimalism with flat pear icon and correct text hierarchy including em-dashes; Image 2's textured pear is less minimalist/Swiss and lacks em-dash format, though still visually appealing. | Pass 2 (B first): Image 2 has correct minimalist Swiss style, cleaner text layout matching exact requested text with em-dashes, and simpler stylised pear. Image 1's pear is textured/realistic, less Swiss-minimalist, and text is split awkwardly without dash format.

tieText rendering — signText in image · 7 v 7.5Pass 1 (A first): Both accurately show text and cobblestone street. Image 1 has richer context, café setting, better lighting and composition.…
PROMPT

A photorealistic wooden café sandwich board on a cobbled street reading exactly "TODAY'S SOUP: TOMATO & BASIL £4.50" in white chalk handwriting.

Gemini · 12.1s · $0.0672
Gemini — Text rendering — sign
GPT Image (in ChatGPT) · 183.1s · $0.2928
GPT Image (in ChatGPT) — Text rendering — sign
JUDGE (blind, position-swapped)

Pass 1 (A first): Both accurately show text and cobblestone street. Image 1 has richer context, café setting, better lighting and composition. Image 2 is clean but simpler background, slightly less photorealistic charm. | Pass 2 (B first): Image 1 has exact matching text including all elements. Image 2 omits some text formatting and appears more cropped, though scene context is richer. Image 1 is cleaner and more accurate to prompt.

GPT Image (in ChatGPT)Flat illustrationIllustration · 6.5 v 8Pass 1 (A first): Image 2 has cleaner flat shapes, no gradients, better cohesive palette matching prompt. Image 1 has gradient shading in…
PROMPT

A flat vector-style illustration of a woman watering a giant potted monstera plant on an apartment balcony, warm sunset palette (coral, amber, deep teal), clean shapes, no gradients, no text.

Gemini · 14.4s · $0.0672
Gemini — Flat illustration
GPT Image (in ChatGPT) · 167.8s · $0.2925
GPT Image (in ChatGPT) — Flat illustration
JUDGE (blind, position-swapped)

Pass 1 (A first): Image 2 has cleaner flat shapes, no gradients, better cohesive palette matching prompt. Image 1 has gradient shading in sky/plants violating 'no gradients' constraint, though both similarly adhere to composition. | Pass 2 (B first): Image 1 has cleaner flat shapes, no gradients, better color palette matching prompt (coral, amber, teal). Image 2 has visible gradients in sky and lighting effects, violating 'no gradients' constraint.

GPT Image (in ChatGPT)Children's bookIllustration · 5.5 v 8Pass 1 (A first): Image1 includes text ('BUS STOP', signage) violating no-text constraint, capping its score. Image2 has no text, warmer…
PROMPT

A children's picture-book illustration of a small red fox in a yellow raincoat sharing an umbrella with a hedgehog at a rainy bus stop, soft watercolour texture, gentle and warm, no text.

Gemini · 8.7s · $0.0685
Gemini — Children's book
GPT Image (in ChatGPT) · 219.1s · $0.2936
GPT Image (in ChatGPT) — Children's book
JUDGE (blind, position-swapped)

Pass 1 (A first): Image1 includes text ('BUS STOP', signage) violating no-text constraint, capping its score. Image2 has no text, warmer composition, gentle watercolor style, and better matches prompt overall. | Pass 2 (B first): Image 1 has yellow raincoat as requested, gentle warm watercolor, no text. Image 2 has text ('BUS STOP') violating no-text constraint and umbrella is yellow not distinct from coat, less adherence overall despite charm.

tiePhotoreal portraitPhotorealism · 8 v 7.5Pass 1 (A first): Image 1 has sharper eyes, better golden hour lighting, richer harbour bokeh, more convincing net-mending action. Image 2 is…
PROMPT

A photorealistic portrait of an elderly fisherman with weathered skin mending a green fishing net at golden hour, 85mm lens look, sharp eyes, softly blurred harbour background.

Gemini · 9.3s · $0.0686
Gemini — Photoreal portrait
GPT Image (in ChatGPT) · 197.3s · $0.2940
GPT Image (in ChatGPT) — Photoreal portrait
JUDGE (blind, position-swapped)

Pass 1 (A first): Image 1 has sharper eyes, better golden hour lighting, richer harbour bokeh, more convincing net-mending action. Image 2 is decent but eyes less sharp, lighting flatter, hands less engaged with net detail. | Pass 2 (B first): Image 1 better matches the 85mm close-up look with softly blurred harbour background and sharp eyes, per prompt. Image 2 has sharp background boats, contradicting the requested blur, though technically strong and detailed.

GPT Image (in ChatGPT)Photoreal scenePhotorealism · 6.5 v 8Pass 1 (A first): Image 2 matches prompt better with empty foreground stools, steam, rain visible, warm lantern light. Image 1 has stools occupied…
PROMPT

A photorealistic interior of a small Tokyo ramen bar at night, steam rising from bowls, warm lantern light, two empty stools in the foreground, rain visible through the window.

Gemini · 12.0s · $0.0672
Gemini — Photoreal scene
GPT Image (in ChatGPT) · 202.1s · $0.2935
GPT Image (in ChatGPT) — Photoreal scene
JUDGE (blind, position-swapped)

Pass 1 (A first): Image 2 matches prompt better with empty foreground stools, steam, rain visible, warm lantern light. Image 1 has stools occupied by people, contradicting 'empty stools' requirement, though technically detailed. Image 2 has cleaner composition and mood. | Pass 2 (B first): Image 1 shows empty stools as requested with steam, rain, lanterns; Image 2 has occupied stools, contradicting 'empty stools' requirement, though it's atmospheric and detailed. Image 1 better adheres to prompt with more coherent composition.

tieInfographicDiagrams · 7 v 7.5Pass 1 (A first): Image 1 has richer, more polished corporate design with icons, arrows, and clear labeling matching the exact text; Image 2 is…
PROMPT

A clean business infographic showing a 4-step process labelled exactly "1. Capture 2. Sort 3. Review 4. Ship" as connected rounded boxes, left to right, corporate blue and white palette, flat design.

Gemini · 8.1s · $0.0689
Gemini — Infographic
GPT Image (in ChatGPT) · 147.7s · $0.2954
GPT Image (in ChatGPT) — Infographic
JUDGE (blind, position-swapped)

Pass 1 (A first): Image 1 has richer, more polished corporate design with icons, arrows, and clear labeling matching the exact text; Image 2 is cleaner but simpler, less detailed, with plain title-less layout. Both meet text requirement, but Image 1 offers stronger visual appeal and completeness. | Pass 2 (B first): Image 1 matches the clean, minimal flat-design request precisely with correct text, simple icons, and clear arrows. Image 2 is more cluttered, adds extra unrequested text/title, deviating from 'clean' minimal brief despite polished visuals.

GPT Image (in ChatGPT)InstructionalDiagrams · 6.5 v 7.5Pass 1 (A first): Both show 3x2 grids with numbered steps in clean line-art. Image 2 has clearer, cleaner progression resembling actual crane…
PROMPT

A simple instructional diagram showing how to fold a paper crane in 6 numbered steps, line-art style, black lines on white, each step clearly separated in a 3x2 grid.

Gemini · 7.2s · $0.0684
Gemini — Instructional
GPT Image (in ChatGPT) · 149.2s · $0.2941
GPT Image (in ChatGPT) — Instructional
JUDGE (blind, position-swapped)

Pass 1 (A first): Both show 3x2 grids with numbered steps in clean line-art. Image 2 has clearer, cleaner progression resembling actual crane folding sequence and ends with recognizable crane. Image 1's step 6 shows two separate cranes, less consistent with single sequence. | Pass 2 (B first): Both show clean 3x2 grids with line-art steps and final crane. Image 1 has clearer, more logical progression and cleaner final crane shape. Image 2's final step shows two cranes which slightly deviates from single instructional flow.

What they cost
Gemini
GPT Image (in ChatGPT)

No verified price claim yet — we show nothing rather than guess. Check the vendor's pricing page.

Measured cost of this exact 10-task run: $0.6803 (Gemini) vs $2.9362 (GPT Image (in ChatGPT)) raw outputs17 Jul 2026
Speed, measured

Median response time across all 10 tasks: 9.3s (Gemini) vs 183.1s (GPT Image (in ChatGPT)).

Receipts — every citation, raw outputs

The full run — prompts, both raw outputs per task, judge reasoning, latencies, token counts and per-call cost — is downloadable as JSON. The task suite is versioned (images-2026-07) and published on the methodology page.

Our verdict — we ran the tasks
Leaning GPT Image (in ChatGPT)

GPT Image (in ChatGPT) took 4 of the 6 tasks that had a clear winner (Gemini 2, GPT Image (in ChatGPT) 4). The judge could pick a winner on 6 of 10 tasks; on the other 4 it could not tell them apart. That is a lean, not a proven win — at this sample size we cannot rule out chance, so we are not calling it decisive.

The audience — what people posted
Gemini
What people say
Not enough data to say
9 of 49 posts gave any opinion — too few to put a number on
GPT Image (in ChatGPT)
What people say
Too few people talking
7 of 47 posts gave any opinion — too few to put a number on
Loudest post this month
How the audience score is measured
Gemini
49 public posts sampled over 90 days; 9 carried a clear view, 40 were announcements or neutral and are excluded from the score. Sources: github (ok), hackernews (ok), reddit (rate-limited), youtube (ok). Classified by anthropic/claude-sonnet-5 under method audience-2026-08-c. Public posts skew negative — people write when something breaks — so this compares like with like rather than rating quality in the absolute. How this is measured
GPT Image (in ChatGPT)
47 public posts sampled over 90 days; 7 carried a clear view, 40 were announcements or neutral and are excluded from the score. Sources: github (ok), hackernews (ok), reddit (partial), youtube (ok). Classified by google/gemini-3.1-pro-preview under method audience-2026-08-c. Public posts skew negative — people write when something breaks — so this compares like with like rather than rating quality in the absolute. How this is measured
Reviewed by Robert Prime
25 years building and selling ecommerce businesses, 15+ exits. Runs MrPrime and trains companies on applied AI.
changelog: 17 Jul 2026 — first published from run #7 · suite images-2026-07
Too close to call 24
raw outputs ↓