Epoch AI's undisclosed OpenAI funding of FrontierMath draws criticism
Epoch AI acknowledged OpenAI funded and had privileged access to FrontierMath's problems and solutions, and had not told contributing mathematicians before the benchmark featured in o3's launch.
- Benchmarks & progress
- Notable
TechCrunch reported that Epoch AI, the research organisation behind the FrontierMath benchmark, had not disclosed to contributing mathematicians that OpenAI funded the benchmark’s creation and had access to its problems and solutions — a relationship Epoch had acknowledged only a month earlier, on the day OpenAI used FrontierMath to help showcase o3.
Stanford PhD student Carina Hong said six mathematicians who had contributed substantially to the benchmark told her they had not known of OpenAI’s involvement, and that most said they were not sure they would have contributed had they known; an Epoch contractor separately called the earlier silence “non-transparent.” Epoch co-founder Tamay Besiroglu said “we made a mistake” on transparency, and that Epoch had been contractually restricted from disclosing the relationship earlier. He said OpenAI had a verbal agreement not to train on the FrontierMath problems, and that Epoch retained a separate holdout set for independent verification of any model OpenAI scored against it.
Days later, Epoch published a fuller account: OpenAI had commissioned 300 of the benchmark’s problems and funded their development, and held rights to the full problem set and its solutions, with one exception — a 50-question holdout for which OpenAI received only the problem statements, not the answers, intended to let outside researchers check OpenAI’s models against unpublished solutions. Epoch also disclosed it could not share FrontierMath’s questions and answers with other parties without OpenAI’s written permission, and committed to disclosing sponsorship and data-access arrangements upfront for future benchmarks.
The episode became a reference case for the tension between commercial funding and benchmark independence: FrontierMath had been positioned as a saturation-resistant, expert-level test of mathematical reasoning, and the disclosure showed one of its funders held privileged access to the material a model would later be scored on.