FrontierMath
FrontierMath is a benchmark of original, unpublished mathematics problems that take expert mathematicians hours to days to solve. It was built to be contamination-proof by construction, since the problems have never been published. It is also the clearest case study in this directory of a benchmark whose credibility problem is about who paid for it rather than how it is built.
| What it measures | Research-level original mathematics requiring hours to days of expert effort, across number theory, analysis, algebraic geometry and more. |
|---|---|
| Built by | Epoch AI, 2024 |
| Format | 338 original, unpublished problems (after the June 2026 v2 correction): 295 in Tiers 1 to 3 plus 43 exceptionally hard Tier 4 problems, each with a verifiable answer |
| Scoring metric | Accuracy (fraction with a correct, automatically verifiable final answer) |
| Status | Saturated |
| Representative top score | 87% (Tiers 1-3) · Claude Fable 5 · read 2026-06 |
| Official leaderboard | epoch.ai/benchmarks/frontiermath |
How FrontierMath works
After a June 2026 correction the set contains 338 original problems: 295 across Tiers 1 to 3 and 43 exceptionally hard Tier 4 problems. Every problem has a single answer that can be verified automatically, which avoids any dependence on a judge model. The problems span number theory, analysis, algebraic geometry and other research fields, and are vetted by expert mathematicians.
History and current status
Epoch AI introduced it in 2024 and kept most of it held back to prevent leakage. A v2 update corrected errors in 42% of problems, which means pre-v2 and post-v2 scores are not comparable at all, a detail routinely dropped when the benchmark is cited. Saturation followed quickly: Epoch reports roughly 87% on Tiers 1 to 3 and 88% on Tier 4 for Claude Fable 5 as of June 2026, with GPT-5.5 Pro statistically tied on the lower tiers.
What the score does not tell you
Epoch AI discloses that FrontierMath was funded by OpenAI, and that OpenAI has exclusive access to a subset of the problems. That is a structural conflict of interest: the organisation being measured helped pay for the test and can see part of it. Epoch deserves credit for disclosing it, but the disclosure does not remove the problem, and it is the reason this benchmark is a standard exhibit in arguments about evaluation independence.
This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, why benchmark saturation makes a high score meaningless.
Who reports FrontierMath, and how to read it
Epoch AI runs the evaluation and publishes results, and labs cite its figures at launch. Two things to check: whether the number is pre-v2 or post-v2, since those are different benchmarks in practice, and which tier it refers to. Epoch’s Open Problems track is now the remaining frontier as the main tiers saturate.
Benchmarks to read alongside this one
PutnamBench
Whether a neural theorem prover can produce a formal, machine-checked proof of an undergraduate competition problem.
AIME 2025
Olympiad-track competition mathematics at the level of the American Invitational Mathematics Examination, used as a high-difficulty LLM eval.
Epoch Capabilities Index
Overall model capability on one continuous scale, stitched together from many benchmarks of differing difficulty.
FrontierMath: frequently asked questions
- What is FrontierMath?
- FrontierMath is a benchmark from Epoch AI of 338 original, unpublished research-level mathematics problems, split into Tiers 1 to 3 (295 problems) and an exceptionally hard Tier 4 (43 problems). Each has a single automatically verifiable answer.
- Who funded FrontierMath?
- OpenAI funded it, and Epoch AI discloses that OpenAI has exclusive access to a subset of the problems. That is a real conflict of interest for a benchmark used to evaluate OpenAI models, and it is why FrontierMath is central to debates about who should build evaluations.
- Are old FrontierMath scores still comparable?
- No. A v2 update corrected errors in 42% of the problems, so scores from before and after that correction measure different sets. Any comparison that spans the v2 boundary is invalid, and a cited FrontierMath figure should state which version it came from.
Get each breakdown before it makes the rounds
You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.
No spam. Unsubscribe anytime.