FrontierMath
FrontierMath is a benchmark of original, unpublished mathematics problems that take expert mathematicians hours to days to solve. It was built to be contamination-proof by construction, since the problems have never been published. It is also the clearest case study in this directory of a benchmark whose credibility problem is about who paid for it rather than how it is built.
| What it measures | Research-level original mathematics requiring hours to days of expert effort, across number theory, analysis, algebraic geometry and more. |
|---|---|
| Built by | Epoch AI, 2024 |
| Format | 338 original, unpublished problems (after the June 2026 v2 correction): 295 in Tiers 1 to 3 plus 43 exceptionally hard Tier 4 problems, each with a verifiable answer |
| Scoring metric | Accuracy (fraction with a correct, automatically verifiable final answer) |
| Status | Saturated |
| Representative top score | 87% (Tiers 1-3) · Claude Fable 5 · read 2026-06 |
| Official leaderboard | epoch.ai/benchmarks/frontiermath |
How FrontierMath works
After a June 2026 correction the set contains 338 original problems: 295 across Tiers 1 to 3 and 43 exceptionally hard Tier 4 problems. Every problem has a single answer that can be verified automatically, which avoids any dependence on a judge model. The problems span number theory, analysis, algebraic geometry and other research fields, and are vetted by expert mathematicians.
History and current status
Epoch AI introduced it in 2024 and kept most of it held back to prevent leakage. A v2 update corrected errors in 42% of problems, which means pre-v2 and post-v2 scores are not comparable at all, a detail routinely dropped when the benchmark is cited. Saturation followed quickly: Epoch reports roughly 87% on Tiers 1 to 3 and 88% on Tier 4 for Claude Fable 5 as of June 2026, with GPT-5.5 Pro statistically tied on the lower tiers.
What the score does not tell you
Epoch AI discloses that FrontierMath was funded by OpenAI, and that OpenAI has exclusive access to a subset of the problems. That is a structural conflict of interest: the organisation being measured helped pay for the test and can see part of it. Epoch deserves credit for disclosing it, but the disclosure does not remove the problem, and it is the reason this benchmark is a standard exhibit in arguments about evaluation independence.
This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, why benchmark saturation makes a high score meaningless.
What FrontierMath scores actually mean
FrontierMath scores need two qualifications before they mean anything. The first is the version: a v2 correction changed 42% of problems, so any figure predating it is not comparable to one after it, and mixing the two produces a fake improvement curve. The second is the tier. Tiers 1 through 3 and the 43 Tier 4 problems are effectively separate benchmarks, and a headline percentage that does not say which tier it covers is unusable. Within Tiers 1 to 3 the leader reads about 87%, which is close enough to the ceiling that the top models are statistically tied. The problems are original and largely unpublished, so contamination is genuinely low, which is what made the rapid saturation notable rather than suspicious.
Who reports FrontierMath, and how to read it
Epoch AI runs the evaluation and publishes results, and labs cite its figures at launch. Two things to check: whether the number is pre-v2 or post-v2, since those are different benchmarks in practice, and which tier it refers to. Epoch’s Open Problems track is now the remaining frontier as the main tiers saturate. Epoch also discloses OpenAI funding and exclusive access to a subset, which belongs beside any OpenAI figure.
When to weight FrontierMath in a model choice
Use FrontierMath only for research-grade mathematical reasoning, which is a narrow question most model selections never ask. If that is the question, quote the tier and the dataset version explicitly and take figures from Epoch AI's own tracker rather than a lab announcement. Weigh the disclosed conflict of interest: Epoch has stated the work was funded by OpenAI, which holds exclusive access to a subset of problems, and that arrangement should be mentioned wherever an OpenAI score on this benchmark is cited. For remaining mathematical headroom, Epoch's Open Problems track has replaced this set as the frontier.
Benchmarks to read alongside this one
PutnamBench
Whether a neural theorem prover can produce a formal, machine-checked proof of an undergraduate competition problem.
AIME 2025
Olympiad-track competition mathematics at the level of the American Invitational Mathematics Examination, used as a high-difficulty LLM eval.
Epoch Capabilities Index
Overall model capability on one continuous scale, stitched together from many benchmarks of differing difficulty.
FrontierMath: frequently asked questions
- What is FrontierMath?
- FrontierMath is a benchmark from Epoch AI of 338 original, unpublished research-level mathematics problems, split into Tiers 1 to 3 (295 problems) and an exceptionally hard Tier 4 (43 problems). Each has a single automatically verifiable answer.
- Who funded FrontierMath?
- OpenAI funded it, and Epoch AI discloses that OpenAI has exclusive access to a subset of the problems. That is a real conflict of interest for a benchmark used to evaluate OpenAI models, and it is why FrontierMath is central to debates about who should build evaluations.
- Are old FrontierMath scores still comparable?
- No. A v2 update corrected errors in 42% of the problems, so scores from before and after that correction measure different sets. Any comparison that spans the v2 boundary is invalid, and a cited FrontierMath figure should state which version it came from.
Sources
- Glazer, Erdil, Besiroglu, Chicharro et al. (2024). FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI. arXiv preprint 2411.04872. arxiv.org/abs/2411.04872
- Epoch AI. FrontierMath project page, including the v2 correction and funding disclosure. epoch.ai/frontiermath
- Epoch AI. FrontierMath benchmark tracker with per-tier results. epoch.ai/benchmarks/frontiermath
Get each breakdown before it makes the rounds
You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.
No spam. Unsubscribe anytime.