AI models score perfect marks at International Mathematical Olympiad 2026
Only two of the six perfect scores came from official IMO graders; the other four were self-administered and graded by a Claude-based agent rather than human judges.
- Benchmarks & progress
- Models & capabilities
- Era-defining
At the International Mathematical Olympiad, held in Shanghai on 15–16 July 2026, six AI systems were reported to have scored full marks — 42 out of 42 points — on the competition’s six problems, a first for the event. Two results, from Huawei’s Celia and Xiaohongshu’s dots-note-3.0, went through the IMO’s own grading process alongside human contestants. The other four — Anthropic’s Claude Fable 5, OpenAI’s GPT-5.6 Sol, Moonshot’s Kimi K3 and Axiom Math’s AxiomProver — sat a self-administered version of the same problems outside the official competition, with a Claude-based agent, rather than human judges, doing the grading.
Coverage of the results emphasised the difference between the two grading tiers: an analysis of the event described the self-graded scores as “strong but not authoritative,” compared with the two officially verified results. Among human contestants, 7 of 666 achieved a perfect score. The performance extended a run of rapid improvement — AI systems reached gold-medal-equivalent scores for the first time only in 2025 — but the analysis also noted that the scores themselves were becoming less informative as a measure: three of the perfect-scoring systems differed in cost by roughly $20 to $51 per run, and adjusting one model’s “effort” setting alone swung its score by 11 points, evidence cited as a sign that headline benchmark numbers were starting to saturate as a way of comparing frontier models.
The result was read alongside a separate, less formally verified claim in the same week that a Claude model had produced a counterexample to the Jacobian Conjecture, an open mathematical problem — both cited in subsequent commentary as evidence that frontier models were reaching or exceeding elite human performance on closed-form mathematical problems, while leaving open how that translated into open-ended research capability.