Gemini with Deep Think reaches gold-medal standard at the 2025 IMO
The IMO itself confirmed the 35/42 score, two days after OpenAI's self-graded claim of the same result; DeepMind said it had waited deliberately for that verification.
- Models & capabilities
- Benchmarks & progress
- Major
Google DeepMind announced that an advanced version of Gemini 2.5 Deep Think had achieved a gold-medal score at the 2025 International Mathematical Olympiad, solving five of six problems for 35 out of 42 points — the gold-medal threshold that year. Unlike DeepMind’s AlphaProof and AlphaGeometry 2 systems, which had reached silver-medal standard at the 2024 IMO by translating problems into the formal language Lean, this model worked end-to-end in natural language, reading the official problem statements directly and producing written proofs within the same 4.5-hour-per-session time limit given to human competitors, using a “parallel thinking” approach that explored multiple solution paths simultaneously.
The distinguishing feature of DeepMind’s announcement was process rather than score: it came with the IMO’s own endorsement. IMO president Gregor Dolinar was quoted confirming the result directly, and DeepMind said it had worked with competition organisers in advance and deliberately withheld its announcement until after that official grading was complete — a contrast the company drew explicitly with OpenAI’s announcement two days earlier, which had used private graders rather than the IMO’s own verification and had been posted before the process organisers requested.
Google said the competition-grade model would first go to trusted testers, including professional mathematicians, before a version reached Google AI Ultra subscribers — with the important caveat that the publicly released Gemini 2.5 Deep Think was a faster, lower-performing variant of the model used in the competition, reaching only bronze-level performance on the same problem set in Google’s internal evaluations. Together with OpenAI’s result, the episode marked the first time general-purpose language models, rather than problem-specific systems, had reached gold-medal standard at the IMO — with the manner of verification, rather than the underlying capability, becoming the point of contention between the two labs.