ETH Zurich launches MathArena live math-competition benchmark
Scoring 30 models on 149 problems from five 2025 competitions, the paper found strong signs older AIME questions were already contaminated and top models scoring below 25% on proof-writing.
- Benchmarks & progress
- Minor
Researchers at ETH Zurich’s SRI Lab published MathArena, a benchmark that evaluates language models on mathematics competitions as soon as new problem sets are released, so that no model can have seen the problems during training. The platform had already been scoring models against 2025 competitions such as AIME and HMMT for several months before the accompanying paper formalised the methodology; the paper itself covered 30 models across 149 problems drawn from five competitions.
The motivation was contamination: static benchmarks built from problems that circulate online eventually leak into training data, inflating scores without reflecting genuine reasoning improvement. The authors reported finding “strong signs of contamination” in AIME 2024, a widely used benchmark, and argued this undermined comparisons built on it. Evaluated instead on freshly released problems, top models still showed what the authors called impressive performance on standard answer-only questions, but the paper’s most striking result concerned proof-writing: on USAMO 2025, a competition requiring full written proofs rather than a single numeric answer, even the best-performing models scored below 25%, exposing a gap between pattern-matching to a final answer and constructing a valid mathematical argument.
MathArena joined a wave of 2025 benchmarks designed around live or rapidly refreshed problem sets, an implicit acknowledgement that fixed evaluation sets had a shelf life measured in months rather than years once frontier labs trained on anything that reached the public web.