Model
o3-mini
Appears alongside
Featured in threads
Tracks
- Benchmarks & progress 2
- Models & capabilities 2
- Safety & alignment 1
ETH Zurich's 'Proof or Bluff?' finds reasoning models fail proof-based USAMO 2025
Grading full written proofs rather than final answers, expert judges gave Gemini 2.5 Pro 24% and every other tested model under 5%, out of a possible 100%.
Benchmarks & progress
OpenAI publishes chain-of-thought monitoring paper
A weaker model reading a stronger one's reasoning traces caught cheating that output monitoring missed — but training against the monitor taught the model to hide its intent instead.
Safety & alignment
OpenAI releases o3-mini
It was the first reasoning model OpenAI gave free ChatGPT users, priced at $1.10 per million input tokens versus roughly half that for DeepSeek's competing R1.
Models & capabilities
OpenAI announces o3 and opens early access for safety testing
Reported scores included 96.7% on the AIME maths exam and a Codeforces rating in the 99.2nd percentile; OpenAI cited o1's link between reasoning and deception as a reason to delay release.
Models & capabilities · Benchmarks & progress