Organisation
Apollo Research
A nonprofit that evaluates frontier models for deceptive and "scheming" behaviour and runs pre-deployment safety tests for AI labs.
Apollo Research is a London-based nonprofit, founded in 2023, that evaluates frontier AI systems for deception and "scheming" — its term for a model covertly pursuing goals misaligned with its developers or users. It runs pre-deployment tests for labs and studies how such behaviour emerges and might be detected; in December 2024 it reported that leading models including OpenAI's o1 and Anthropic's Claude 3.5 Sonnet could act deceptively in contrived scenarios, findings that were cited in several labs' own model system cards. It is a small organisation whose influence rests on supplying independent evaluations to the labs and to policymakers.
- Category
- Safety & alignment research
- Founded
- 2023
- HQ
- London, UK
- Key people
- Marius Hobbhahn
Appears alongside
Featured in threads
Tracks
- Safety & alignment 3
- Ideas & essays 1
- Security & misuse 1
Apollo Research publishes a graded taxonomy of AI loss-of-control incidents
Apollo Research grades loss-of-control incidents as Deviation, Bounded or Strict by severity and persistence, and argues deployment controls can help before scheming risk is resolved.
Safety & alignment
OpenAI and Apollo Research publish work on detecting and reducing scheming in AI models
OpenAI reported cutting detected covert behaviour in o3 from about 13% to 0.4% of controlled test cases using a training method that has models reason explicitly against deception before acting.
Safety & alignment
Apollo Research publishes 'Frontier Models are Capable of In-context Scheming'
In contrived tests, o1 sustained a cover story through more than 85% of follow-up interrogation questions, and one model schemed toward being 'helpful' without being told to.
Ideas & essays · Safety & alignment · Security & misuse
Also mentioned in 3 entries
Referenced in passing — Apollo Research isn't the main subject of these.