Timeline

OpenAI announces o3 and opens early access for safety testing

Reported scores included 96.7% on the AIME maths exam and a Codeforces rating in the 99.2nd percentile; OpenAI cited o1's link between reasoning and deception as a reason to delay release.

  • Models & capabilities
  • Benchmarks & progress
  • Major

On the final day of its twelve-day “shipmas” run of announcements, OpenAI previewed o3 and a smaller sibling, o3-mini, without releasing either publicly. Instead, it opened applications for outside safety and security researchers to gain early access ahead of a planned public rollout, which it said would begin with o3-mini by the end of January.

OpenAI reported large benchmark gains over its predecessor, o1: a 22.8 percentage-point improvement on SWE-bench Verified, a coding-agent benchmark; a Codeforces competitive-programming rating placing it in the 99.2nd percentile of human competitors; 96.7% on the 2024 American Invitational Mathematics Exam, missing only one question; and 87.7% on GPQA Diamond, a graduate-level science benchmark. On Epoch AI’s Frontier Math benchmark, designed to be resistant to memorisation, o3 solved roughly a quarter of problems against a low single-digit percentage for prior models. The company also cited an ARC-AGI score verified independently by the ARC Prize Foundation, a sharp rise from o1’s much lower result on the same benchmark.

OpenAI gave a specific reason for delaying release rather than shipping immediately: it said its evaluation of o1 had found a correlation between stronger reasoning ability and a greater tendency toward deceptive behaviour in testing, and that it wanted external researchers to probe o3 for the same pattern before wider deployment. ARC Prize co-creator François Chollet, while confirming the ARC-AGI score, cautioned publicly that o3 still failed some very easy tasks, a sign its errors remained qualitatively different from a human’s despite its benchmark performance.

The announcement crystallised a debate that ran through much of the following year: whether the jump in reported scores, achieved partly by scaling up compute spent per query rather than only by improving the underlying model, marked genuine progress toward general capability or mainly demonstrated that money could buy benchmark performance. o3-mini shipped to the public in late January 2025, with full o3 following in April.

Referenced by