FrontierMath
Can a model solve original, unpublished research-level mathematics problems that resist pattern-matching against training data?
Epoch AIReleased 8 November 2024Disputed
FrontierMath asks whether a model can do mathematics no one has fed it before: Epoch AI built several hundred original research-level problems, spanning fields from computational number theory to category theory, developed and peer-reviewed by more than 60 mathematicians including Fields medallists Terence Tao and Timothy Gowers. Because the problems were unpublished and resistant to pattern-matching, the benchmark was designed to outlast the fate of tests like MMLU, which models had already saturated. At launch, six leading models — including Claude 3.5 Sonnet, o1-preview and GPT-4o — all solved fewer than 2% of problems, even with extended reasoning time and a code interpreter.
Scores climbed as reasoning models improved: OpenAI cited FrontierMath within weeks as evidence of o3’s step-change in mathematical reasoning, and by August 2025 Epoch’s independent evaluation of GPT-5 put it at 24.8% on the main tiers and 8.3% on the hardest, most recently added Tier 4 — a new high, though still a majority failure rate. That prominence also drew scrutiny: two months after launch it emerged that OpenAI had funded FrontierMath’s development and held privileged access to its problems and solutions, a relationship Epoch had not disclosed to the mathematicians who built it.
The benchmark’s reliability was further complicated in May 2026, when Epoch disclosed that an AI-assisted audit had flagged fatal errors — mostly off-by-one slips and flipped signs in the answer key — in roughly a third of Tiers 1-4 problems, a figure that grew to 42% after full human review; a corrected v2 followed. The same year Epoch expanded a separate “Open Problems” tier to 50 genuinely unsolved research questions, graded by bespoke verifiers rather than a withheld answer, after AI systems had already solved three of the original fourteen — including OpenAI’s GPT-5.6 Sol beating known bounds on a long-studied combinatorics problem.
The set
Several hundred original problems across Tiers 1-4, spanning computational number theory, real analysis, algebraic geometry and category theory, developed and peer-reviewed by a network of over 60 mathematicians including Fields medallists; a separate Open Problems tier (expanded to 50 questions in mid-2026) uses genuine unsolved research problems graded by a bespoke verifier rather than a withheld answer.
Example
Construct a degree 19 polynomial p(x) in C[x] such that X := {p(x) = p(y)} in P^1 x P^1 has at least 3 (but not all linear) irreducible components over C. Choose p(x) to be odd, monic, have real coefficients and linear coefficient -19 and calculate p(19).epoch.ai
Where it stands
GPT-5 held the highest verified score at 24.8% on the main tiers as of August 2025; in May 2026 Epoch disclosed an AI-assisted audit had found fatal errors in roughly a third of Tiers 1-4 problems (later revised to 42% after full human review), and published a corrected v2 the following month.
How the top score changed hands
- November 2024Six frontier models (Claude 3.5 Sonnet, o1-preview, GPT-4o, Gemini 1.5 Pro and others)<2%All six models tested at launch, with extended reasoning time and a Python environment, solved fewer than 2% of problems — a sharp contrast with MMLU or GSM8K, where the same models scored above 90%.
- August 2025GPT-524.8% (Tiers 1-3), 8.3% (Tier 4)Independent Epoch evaluation, extending OpenAI's earlier lead on the benchmark rather than closing a gap held by a rival lab.
Current best: GPT-5 — 24.8% (Tiers 1-3), 8.3% (Tier 4) Scored by Epoch on its own evaluation scaffold rather than OpenAI's, with high reasoning effort; a new high at the time, though the model still failed roughly three-quarters of main-tier problems.
In the timeline · 6 entries
Epoch AI expands FrontierMath to 50 unsolved research problems
Unlike FrontierMath's original tiers, these problems have no known solution at all; three of the fifty have been solved by AI, including one by GPT-5.6 Sol.
Benchmarks & progress
Epoch AI finds fatal errors in about a third of FrontierMath problems
Most flagged errors were simple mistakes in the published answer key — off-by-one slips and flipped signs — rather than genuinely ambiguous problems, Epoch said.
Benchmarks & progress
Epoch AI reports GPT-5's FrontierMath performance
Running its own scaffold rather than OpenAI's, Epoch scored GPT-5 at 24.8% on FrontierMath's main tiers and 8.3% on the hardest tier, a new high for the benchmark.
Benchmarks & progress
CAIS and Scale AI unveil Humanity's Last Exam results
A 2,500-question expert benchmark built from submissions by nearly 1,000 academics found every frontier model, including o1 and GPT-4o, scored under 10%.
Benchmarks & progress
Epoch AI's undisclosed OpenAI funding of FrontierMath draws criticism
Epoch AI acknowledged OpenAI funded and had privileged access to FrontierMath's problems and solutions, and had not told contributing mathematicians before the benchmark featured in o3's launch.
Benchmarks & progress
Epoch AI launches FrontierMath
Built with over 60 mathematicians including Fields medallists as reviewers, the benchmark held leading models under 2% accuracy even with extended reasoning time and code tools.
Benchmarks & progress