Capital & Compute
Benchmark· Mathematics· Checked 2026-07-26

FrontierMath

FrontierMath is a benchmark of original, unpublished mathematics problems that take expert mathematicians hours to days to solve. It was built to be contamination-proof by construction, since the problems have never been published. It is also the clearest case study in this directory of a benchmark whose credibility problem is about who paid for it rather than how it is built.

Key facts about the FrontierMath benchmark
What it measuresResearch-level original mathematics requiring hours to days of expert effort, across number theory, analysis, algebraic geometry and more.
Built byEpoch AI, 2024
Format338 original, unpublished problems (after the June 2026 v2 correction): 295 in Tiers 1 to 3 plus 43 exceptionally hard Tier 4 problems, each with a verifiable answer
Scoring metricAccuracy (fraction with a correct, automatically verifiable final answer)
StatusSaturated
Representative top score87% (Tiers 1-3) · Claude Fable 5 · read 2026-06
Official leaderboardepoch.ai/benchmarks/frontiermath

How FrontierMath works

After a June 2026 correction the set contains 338 original problems: 295 across Tiers 1 to 3 and 43 exceptionally hard Tier 4 problems. Every problem has a single answer that can be verified automatically, which avoids any dependence on a judge model. The problems span number theory, analysis, algebraic geometry and other research fields, and are vetted by expert mathematicians.

History and current status

Epoch AI introduced it in 2024 and kept most of it held back to prevent leakage. A v2 update corrected errors in 42% of problems, which means pre-v2 and post-v2 scores are not comparable at all, a detail routinely dropped when the benchmark is cited. Saturation followed quickly: Epoch reports roughly 87% on Tiers 1 to 3 and 88% on Tier 4 for Claude Fable 5 as of June 2026, with GPT-5.5 Pro statistically tied on the lower tiers.

What the score does not tell you

Epoch AI discloses that FrontierMath was funded by OpenAI, and that OpenAI has exclusive access to a subset of the problems. That is a structural conflict of interest: the organisation being measured helped pay for the test and can see part of it. Epoch deserves credit for disclosing it, but the disclosure does not remove the problem, and it is the reason this benchmark is a standard exhibit in arguments about evaluation independence.

This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, why benchmark saturation makes a high score meaningless.

Who reports FrontierMath, and how to read it

Epoch AI runs the evaluation and publishes results, and labs cite its figures at launch. Two things to check: whether the number is pre-v2 or post-v2, since those are different benchmarks in practice, and which tier it refers to. Epoch’s Open Problems track is now the remaining frontier as the main tiers saturate.

2024
First released
Epoch AI
Saturated
Status today
As of July 27, 2026
87% (Tiers 1-3)
Representative top score
Read 2026-06

Benchmarks to read alongside this one

FrontierMath: frequently asked questions

What is FrontierMath?
FrontierMath is a benchmark from Epoch AI of 338 original, unpublished research-level mathematics problems, split into Tiers 1 to 3 (295 problems) and an exceptionally hard Tier 4 (43 problems). Each has a single automatically verifiable answer.
Who funded FrontierMath?
OpenAI funded it, and Epoch AI discloses that OpenAI has exclusive access to a subset of the problems. That is a real conflict of interest for a benchmark used to evaluate OpenAI models, and it is why FrontierMath is central to debates about who should build evaluations.
Are old FrontierMath scores still comparable?
No. A v2 update corrected errors in 42% of the problems, so scores from before and after that correction measure different sets. Any comparison that spans the v2 boundary is invalid, and a cited FrontierMath figure should state which version it came from.

Get each breakdown before it makes the rounds

You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.

No spam. Unsubscribe anytime.

← All 116 benchmarks in the directory