Benchmarks · Mathematics

FrontierMath

Can a model solve original, unpublished research-level mathematics problems that resist pattern-matching against training data?

Epoch AIReleased 8 November 2024Disputed

FrontierMath asks whether a model can do mathematics no one has fed it before: Epoch AI built several hundred original research-level problems, spanning fields from computational number theory to category theory, developed and peer-reviewed by more than 60 mathematicians including Fields medallists Terence Tao and Timothy Gowers. Because the problems were unpublished and resistant to pattern-matching, the benchmark was designed to outlast the fate of tests like MMLU, which models had already saturated. At launch, six leading models — including Claude 3.5 Sonnet, o1-preview and GPT-4o — all solved fewer than 2% of problems, even with extended reasoning time and a code interpreter.

Scores climbed as reasoning models improved: OpenAI cited FrontierMath within weeks as evidence of o3’s step-change in mathematical reasoning, and by August 2025 Epoch’s independent evaluation of GPT-5 put it at 24.8% on the main tiers and 8.3% on the hardest, most recently added Tier 4 — a new high, though still a majority failure rate. By December, Epoch had Gemini 3 Pro at 38% and GPT-5.2 at 40.3% on Tiers 1-3, the model using the benchmark’s native Python tool. That prominence also drew scrutiny: two months after launch it emerged that OpenAI had funded FrontierMath’s development and held privileged access to its problems and solutions, a relationship Epoch had not disclosed to the mathematicians who built it.

The benchmark’s reliability was further complicated in May 2026, when Epoch disclosed that an AI-assisted audit had flagged fatal errors — mostly off-by-one slips and flipped signs in the answer key — in roughly a third of Tiers 1-4 problems, a figure that grew to 42% after full human review; a corrected v2 followed. The same year Epoch expanded a separate “Open Problems” tier to 50 genuinely unsolved research questions, graded by bespoke verifiers rather than a withheld answer, after AI systems had already solved three of the original fourteen — including OpenAI’s GPT-5.6 Sol beating known bounds on a long-studied combinatorics problem.

The set

Several hundred original problems across Tiers 1-4, spanning computational number theory, real analysis, algebraic geometry and category theory, developed and peer-reviewed by a network of over 60 mathematicians including Fields medallists; a separate Open Problems tier (expanded to 50 questions in mid-2026) uses genuine unsolved research problems graded by a bespoke verifier rather than a withheld answer.

Example

Construct a degree 19 polynomial p(x) in C[x] such that X := {p(x) = p(y)} in P^1 x P^1 has at least 3 (but not all linear) irreducible components over C. Choose p(x) to be odd, monic, have real coefficients and linear coefficient -19 and calculate p(19).epoch.ai

Where it stands

GPT-5 held the highest verified score at 24.8% on the main tiers in August 2025; by December, Gemini 3 Pro had reached 38% and GPT-5.2 40.3% on Tiers 1-3. In May 2026 Epoch disclosed an AI-assisted audit had found fatal errors in roughly a third of Tiers 1-4 problems (later revised to 42% after full human review), and published a corrected v2 the following month, so scores from before mid-2026 are against the original problem set.

How the top score changed hands

In the timeline · 7 entries

More mathematics benchmarks