If you write code
These compare things you reach through an API key or a developer account, and the tasks behind them are programming tasks. Same protocol as everything else on the site: identical prompts through one harness, three judges from three labs, position-swapped, every raw output downloadable. Looking for a chatbot instead? →
Coding
Probably GitHub Copilot
Leans · not proven yet
It was also the faster and the cheaper of the two.
5 of 18 tasks had a clear winner
tested 7 Aug 2026
see why — full evidence →
Best free
Probably Gemini 3.5 Flash
Leans · not proven yet
Worth knowing: DeepSeek V4 Flash costs noticeably less.
11 of 30 tasks had a clear winner
tested 7 Aug 2026
see why — full evidence →
Best-value API
Either one works
Too close
Pick GPT-5.6 Terra for speed, or DeepSeek V4 Pro to spend less.
11 of 30 tasks had a clear winner
tested 7 Aug 2026
see why — full evidence →
Model and API changes
- GLM 5.2 output price down 59% — output price 1.9316 → 0.792 over 29 changes8 Aug 2026
- GLM 5.2 input price down 59% — input price 0.6146 → 0.252 over 29 changes8 Aug 2026
- Claude Opus 4.6 slipped to 5 on the leaderboard, from 4 — LMArena leaderboard (CC-BY-4.0)7 Aug 2026
- Qwen3.7 Max slipped to 16 on the leaderboard, from 14 — LMArena leaderboard (CC-BY-4.0)7 Aug 2026
- Gemini 3.5 Flash slipped to 10 on the leaderboard, from 9 — LMArena leaderboard (CC-BY-4.0)4 Aug 2026
- Gemini 3.1 Pro Preview slipped to 12 on the leaderboard, from 11 — LMArena leaderboard (CC-BY-4.0)7 Aug 2026
- Grok 4.5 leaderboard place up 9% — leaderboard place 34 → 37 over 2 changes7 Aug 2026
- Claude Opus 4.8 leaderboard place up 6% — leaderboard place 33 → 35 over 2 changes7 Aug 2026
- DeepSeek V4 Pro leaderboard place up 5% — leaderboard place 37 → 39 over 2 changes7 Aug 2026
- GPT-5.5 leaderboard place up 4% — leaderboard place 23 → 24 over 2 changes7 Aug 2026