Organisation · nonprofit · UK

Apollo Research

A nonprofit that evaluates frontier models for deceptive and "scheming" behaviour and runs pre-deployment safety tests for AI labs.

3 entries · December 2024 – November 2025

Apollo Research is a London-based nonprofit, founded in 2023, that evaluates frontier AI systems for deception and "scheming" — its term for a model covertly pursuing goals misaligned with its developers or users. It runs pre-deployment tests for labs and studies how such behaviour emerges and might be detected; in December 2024 it reported that leading models including OpenAI's o1 and Anthropic's Claude 3.5 Sonnet could act deceptively in contrived scenarios, findings that were cited in several labs' own model system cards. It is a small organisation whose influence rests on supplying independent evaluations to the labs and to policymakers.

Category
Safety & alignment research
Founded
2023
HQ
London, UK
Key people
Marius Hobbhahn

Tracks

  • Safety & alignment 3
  • Ideas & essays 1
  • Security & misuse 1

Also mentioned in 3 entries

Referenced in passing — Apollo Research isn't the main subject of these.