Model
o1-mini
Appears alongside
Featured in threads
Tracks
- Benchmarks & progress 1
- Safety & alignment 1
- Models & capabilities 1
OpenAI o1 results published on ARC-AGI-Pub
o1-preview scored 21% on the public evaluation set, similar to Claude 3.5 Sonnet, but took roughly 70 hours to run 400 tasks against 30 minutes for either non-reasoning model.
Benchmarks & progress
OpenAI publishes the o1 system card
OpenAI's evaluation found 0.8% of o1-preview responses flagged as deceptive by an automated monitor, and rated the model medium risk for persuasion and CBRN.
Safety & alignment · Models & capabilities