Benchmarks · Mathematics

MathArena

also: MathArena.ai

Can a model solve maths competition problems released after its training cutoff, so a score reflects reasoning rather than memorised answers or leaked solutions?

ETH Zurich SRI Lab & INSAITReleased May 2025Live

MathArena exists to solve a problem that static maths benchmarks cannot: once a problem set is public, it eventually leaks into training data, and a rising score stops proving anything about reasoning. Researchers at ETH Zurich’s SRI Lab and INSAIT built it as a live platform instead of a fixed test, scoring models on competition mathematics within days of each new sitting — AIME and HMMT among the final-answer contests, USAMO, IMO and the Putnam among those requiring a full written proof, which MathArena grades with qualified human judges rather than an automated checker or another model.

The launch paper made the case for the approach directly, reporting “strong signs of contamination” in AIME 2024, a benchmark by then cited in nearly every frontier release. Its proof-based results were more striking still: on USAMO 2025, graded on the actual logical validity of a written argument rather than a final number, even the best-performing models scored under 25%. A companion study by the same group, Proof or Bluff?, reached a similar verdict on a different olympiad and found LLM graders inflated scores by up to 20 times against expert human judges.

MathArena has since grown well past its original five-competition, 30-model launch paper, and its live overall ranking — a composite “expected performance” figure across the current competition roster rather than one number — has become one of the standard reference points cited alongside AIME and FrontierMath in coverage of frontier model releases. The gap it keeps re-exposing, between fluent final answers and rigorous proof, has remained wider than headline accuracy figures suggest.

The set

A continuously updated evaluation platform, not a fixed problem set: models are scored on competitions as soon as each year's problems are published, across final-answer contests (AIME, HMMT, CMIMC) and proof-based ones (USAMO, IMO, Putnam, Miklós Schweitzer). Final-answer problems are graded automatically; proofs are graded by qualified human judges rather than an LLM grader, which the project's companion study found could otherwise inflate scores by up to 20-fold.

Example

USAMO 2025, Problem 1, verbatim: Let k and d be positive integers. Prove that there exists a positive integer N such that for every odd integer n > N, the digits in the base-2n representation of n^k are all greater than d.huggingface.co

Where it stands

Overall model rankings blend results across many recent competitions rather than one fixed test; a companion ETH Zurich study using the same methodology found top models scored under 25% on full written proofs even where their final-answer accuracy looked strong.

How the top score changed hands

  1. May 202530 evaluated models (GPT, Claude, Gemini, DeepSeek and others)USAMO 2025 proof score below 25%Launch paper covered 149 problems across five 2025 competitions; separately found strong signs AIME 2024 was already contaminated in training data.
  2. July 2026Claude Opus 5 (max)84.4% ± 2.8%

Current best: Claude Opus 5 (max) — 84.4% ± 2.8% (overall expected performance) MathArena's own composite ranking across its current competition roster, not a single fixed test; GPT-5.6-Sol (max) followed at 79.7%, GPT-5.5 (xhigh) at 77.8%, and Kimi K3 (Think) led open-weight models at 69.7%.

In the timeline · 1 entry

More mathematics benchmarks