Capital & Compute

Models & benchmarks

Leaderboards compress messy reality into a single number, then labs quote whichever number flatters the release. Suites saturate, contamination inflates scores, and the same model under different scaffolding can move ten points. What matters is which benchmarks still separate frontier models, and what each one refuses to measure.

In this world
67 analyses
Field tool
Explore the value leaderboard ↗

Keep diving.

The featured guide stays above; the stream below moves as new analysis is published.

What this topic covers

Model releases, capability comparisons and rigorous analysis of the benchmarks used to rank them.

This topic collects 67 analyses, and they are written to be read together rather than one at a time: the featured guide sets out the shape of the problem, and the pieces below work through the individual numbers, tradeoffs and edge cases behind it. Every figure is attributed to a primary source at the point it is used and carries the date it was verified, because prices and benchmark results in this area go stale in weeks rather than years. Where a number is modeled rather than measured, the assumptions are stated so the arithmetic can be checked.

The topics on this site overlap by design. This one runs into Coding agents and Local AI & hardware, and a question that starts in one usually ends in another: a pricing decision turns into a hardware decision, a benchmark result turns into a cost question. Follow the links inside the posts rather than treating these archives as separate shelves.

Explore the value leaderboardTrack new and upcoming modelsLook up a benchmark