OpenAI and DeepMind reach gold-medal standard at the IMO
OpenAI announced its result on X the day the student competition ended, using its own hired graders rather than the IMO's official verification, drawing criticism from Google.
- Benchmarks & progress
- Models & capabilities
- Major
OpenAI announced on X that an experimental general-purpose reasoning model had scored 35 out of 42 points on the 2025 International Mathematical Olympiad problem set — equal to that year’s gold-medal cutoff — solving five of the six problems under the same conditions given to human contestants: two 4.5-hour sessions, no tools or internet access, and natural-language proofs written from the official problem statements. OpenAI said three former IMO medalists it hired independently graded each submission and reached the scores by unanimous consensus. Researcher Alexander Wei said the model was a separate experimental system, distinct from GPT-5, which the company was preparing to release separately.
The announcement drew criticism over its timing rather than its substance. The IMO organisers had asked participating AI labs to hold results until after the official grading process concluded and the actual student competitors had received their medals and public recognition. OpenAI posted its result on the Saturday morning immediately after the student competition ended, using its own hired graders rather than IMO-certified evaluation. Google DeepMind, which announced its own gold-medal result two days later after obtaining official IMO grading, made the contrast explicit: CEO Demis Hassabis said DeepMind had “respected the IMO Board’s original request that all AI labs share their results only after the official results had been verified by independent experts,” while DeepMind’s Thang Luong argued that any claim not graded under that process “could not make any claim about gold-medal level” performance.
The dispute mattered because it separated two questions that got conflated in headlines: whether a model could solve IMO-level problems (both labs’ models plausibly could) and whether a specific score had been independently verified against the standard used to certify human gold medalists (only Google’s had, at the time of its announcement). The model OpenAI used was not released to the public.