Timeline

Epoch AI finds fatal errors in about a third of FrontierMath problems

Most flagged errors were simple mistakes in the published answer key — off-by-one slips and flipped signs — rather than genuinely ambiguous problems, Epoch said.

  • Benchmarks & progress
  • Notable

Epoch AI disclosed that an AI-assisted review of FrontierMath, the mathematics benchmark it had positioned as a gold-standard test of frontier models’ reasoning, had flagged fatal errors in roughly a third of the problems across Tiers 1-4. Epoch said it believed most of the flags were valid.

We are conducting an AI-assisted review of FrontierMath: Tiers 1-4. This has flagged fatal errors in about a third of problems, and we believe most of these flags to be valid. We will release updated scores on a corrected dataset after completing a thorough human review.

Epoch AI, public statement

Epoch characterised most of the errors as mundane: off-by-one mistakes and flipped signs made while the benchmark’s authors extracted a final numeric answer, rather than problems that were conceptually wrong. A smaller number of problem statements were fatally ambiguous. The company said it would complete a full human review of every flagged problem before publishing corrected scores.

The finding was significant less for its arithmetic than for what it implied about benchmark trust: FrontierMath, which OpenAI had funded and used to claim progress on frontier mathematical reasoning with its o3 model, was already under scrutiny over that funding relationship, and a substantial defect rate in the answer key meant reported scores on the affected problems could not be taken at face value until the dataset was corrected. Epoch published FrontierMath v2 the following month with errors addressed across 42% of problems, a figure larger than the initial one-third estimate once the full human review was complete.