Are AI Benchmarks Reliable? How the Scores Get Gamed
AI benchmark scores get gamed by contamination, saturation, and cheating. The 2026 receipts, a trust scorecard, and how to read past any leaderboard.
Capability under pressure
Leaderboards compress messy reality into one number. Learn what survives outside the test.
Topic archive · page 6
Model releases, capability comparisons and rigorous analysis of the benchmarks used to rank them.
AI benchmark scores get gamed by contamination, saturation, and cheating. The 2026 receipts, a trust scorecard, and how to read past any leaderboard.
Cohere North Mini Code is free on the API and open-weight. Here is what it really costs per task once you self-host it on a single H100.
GPT-5.6 launched June 26 as Sol at $5/$30, half Claude Fable 5 at $10/$50. Which flagship costs less per task, now that Fable 5 is back from suspension.
Google ended free Gemini CLI access on June 18, 2026. The best alternatives (Claude Code, Aider, OpenCode, Antigravity) ranked by real cost per task.
Claude Fable 5 costs exactly 2x Opus 4.8 per token: $10/$50 vs $5/$25. Whether it is cheaper per task depends on loop count, not the sticker rate.
Gemini 3.5 Flash lists at $1.50/$9.00 per million tokens: about 2x Haiku per token, yet roughly 3x cheaper per task than GPT-5.5 on agentic coding work.
Microsoft calls MAI-Code-1-Flash its cheapest coding model. In Copilot it bills like Claude Haiku 4.5 (0.33x); its token edge holds only on easy benchmarks.
Qwen 3.7 Max lists at half Claude Opus 4.8 and runs the same eval suite for a third of the cost. The catch is not hidden cost. It is what the price buys.
Codex leads Claude Code 83.4% to 78.9% on Terminal-Bench 2.1, but the prices match and the cheaper agent flips with your harness. The honest scorecard.