Capital & Compute

FrontierMath

Benchmark· Mathematics· Checked 2026-07-26

FrontierMath is a benchmark of original, unpublished mathematics problems that take expert mathematicians hours to days to solve. It was built to be contamination-proof by construction, since the problems have never been published. It is also the clearest case study in this directory of a benchmark whose credibility problem is about who paid for it rather than how it is built.

Key facts about the FrontierMath benchmark
What it measuresResearch-level original mathematics requiring hours to days of expert effort, across number theory, analysis, algebraic geometry and more.
Built byEpoch AI, 2024
Format338 original, unpublished problems (after the June 2026 v2 correction): 295 in Tiers 1 to 3 plus 43 exceptionally hard Tier 4 problems, each with a verifiable answer
Scoring metricAccuracy (fraction with a correct, automatically verifiable final answer)
StatusSaturated
Representative top score87% (Tiers 1-3) · Claude Fable 5 · read 2026-06
Official leaderboardepoch.ai/benchmarks/frontiermath

How FrontierMath works

After a June 2026 correction the set contains 338 original problems: 295 across Tiers 1 to 3 and 43 exceptionally hard Tier 4 problems. Every problem has a single answer that can be verified automatically, which avoids any dependence on a judge model. The problems span number theory, analysis, algebraic geometry and other research fields, and are vetted by expert mathematicians.

History and current status

Epoch AI introduced it in 2024 and kept most of it held back to prevent leakage. A v2 update corrected errors in 42% of problems, which means pre-v2 and post-v2 scores are not comparable at all, a detail routinely dropped when the benchmark is cited. Saturation followed quickly: Epoch reports roughly 87% on Tiers 1 to 3 and 88% on Tier 4 for Claude Fable 5 as of June 2026, with GPT-5.5 Pro statistically tied on the lower tiers.

What the score does not tell you

Epoch AI discloses that FrontierMath was funded by OpenAI, and that OpenAI has exclusive access to a subset of the problems. That is a structural conflict of interest: the organisation being measured helped pay for the test and can see part of it. Epoch deserves credit for disclosing it, but the disclosure does not remove the problem, and it is the reason this benchmark is a standard exhibit in arguments about evaluation independence.

This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, why benchmark saturation makes a high score meaningless.

What FrontierMath scores actually mean

FrontierMath scores need two qualifications before they mean anything. The first is the version: a v2 correction changed 42% of problems, so any figure predating it is not comparable to one after it, and mixing the two produces a fake improvement curve. The second is the tier. Tiers 1 through 3 and the 43 Tier 4 problems are effectively separate benchmarks, and a headline percentage that does not say which tier it covers is unusable. Within Tiers 1 to 3 the leader reads about 87%, which is close enough to the ceiling that the top models are statistically tied. The problems are original and largely unpublished, so contamination is genuinely low, which is what made the rapid saturation notable rather than suspicious.

Who reports FrontierMath, and how to read it

Epoch AI runs the evaluation and publishes results, and labs cite its figures at launch. Two things to check: whether the number is pre-v2 or post-v2, since those are different benchmarks in practice, and which tier it refers to. Epoch’s Open Problems track is now the remaining frontier as the main tiers saturate. Epoch also discloses OpenAI funding and exclusive access to a subset, which belongs beside any OpenAI figure.

When to weight FrontierMath in a model choice

Use FrontierMath only for research-grade mathematical reasoning, which is a narrow question most model selections never ask. If that is the question, quote the tier and the dataset version explicitly and take figures from Epoch AI's own tracker rather than a lab announcement. Weigh the disclosed conflict of interest: Epoch has stated the work was funded by OpenAI, which holds exclusive access to a subset of problems, and that arrangement should be mentioned wherever an OpenAI score on this benchmark is cited. For remaining mathematical headroom, Epoch's Open Problems track has replaced this set as the frontier.

2024
First released
Epoch AI
Saturated
Status today
As of September 4, 2026
87% (Tiers 1-3)
Representative top score
Read 2026-06

Benchmarks to read alongside this one

FrontierMath: frequently asked questions

What is FrontierMath?
FrontierMath is a benchmark from Epoch AI of 338 original, unpublished research-level mathematics problems, split into Tiers 1 to 3 (295 problems) and an exceptionally hard Tier 4 (43 problems). Each has a single automatically verifiable answer.
Who funded FrontierMath?
OpenAI funded it, and Epoch AI discloses that OpenAI has exclusive access to a subset of the problems. That is a real conflict of interest for a benchmark used to evaluate OpenAI models, and it is why FrontierMath is central to debates about who should build evaluations.
Are old FrontierMath scores still comparable?
No. A v2 update corrected errors in 42% of the problems, so scores from before and after that correction measure different sets. Any comparison that spans the v2 boundary is invalid, and a cited FrontierMath figure should state which version it came from.

Sources

  • Glazer, Erdil, Besiroglu, Chicharro et al. (2024). FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI. arXiv preprint 2411.04872. arxiv.org/abs/2411.04872
  • Epoch AI. FrontierMath project page, including the v2 correction and funding disclosure. epoch.ai/frontiermath
  • Epoch AI. FrontierMath benchmark tracker with per-tier results. epoch.ai/benchmarks/frontiermath

Get each breakdown before it makes the rounds

You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.

No spam. Unsubscribe anytime.

← All 116 benchmarks in the directory