Model
o1-preview
Featured in threads
Tracks
- Safety & alignment 2
- Labs & people 1
- Benchmarks & progress 1
- Models & capabilities 1
OpenAI updates safety and security practices after o1 release
The Safety and Security Committee became an independent board oversight body chaired by CMU professor Zico Kolter, with authority to delay model releases.
Safety & alignment · Labs & people
OpenAI o1 results published on ARC-AGI-Pub
o1-preview scored 21% on the public evaluation set, similar to Claude 3.5 Sonnet, but took roughly 70 hours to run 400 tasks against 30 minutes for either non-reasoning model.
Benchmarks & progress
OpenAI publishes the o1 system card
OpenAI's evaluation found 0.8% of o1-preview responses flagged as deceptive by an automated monitor, and rated the model medium risk for persuasion and CBRN.
Safety & alignment · Models & capabilities