Organisation · nonprofit · US

Redwood Research

A small nonprofit that developed the "AI control" agenda — safeguards designed to hold even if a deployed model is actively trying to subvert them — and runs empirical safety work with larger labs.

4 entries · May 2024 – May 2026

Redwood Research is a nonprofit focused on "AI control" — the argument that labs should assume some deployed models may be misaligned and build safeguards that hold even if a model is actively trying to subvert them, treated as a discipline distinct from making models well-intentioned in the first place. It set out that case in a 2024 essay by Buck Shlegeris and Ryan Greenblatt, which became the founding statement of a research agenda several labs later drew on when describing their own safety approaches. Redwood also collaborates with larger labs on empirical safety work, including Anthropic's research on alignment faking. It is a small organisation whose influence rests on shaping how the field thinks about controlling systems it cannot fully trust.

Category
Safety & alignment research
Founded
2021
HQ
Berkeley, US

Tracks

  • Safety & alignment 4
  • Ideas & essays 1