Timeline

Redwood Research publishes 'The case for ensuring that powerful AIs are controlled'

Argued labs should assume some deployed models may be misaligned and build restrictions that hold even if a model actively tries to subvert them, distinct from alignment itself.

  • Ideas & essays
  • Safety & alignment
  • Notable

Redwood Research’s Buck Shlegeris and Ryan Greenblatt published an essay arguing that AI labs should pursue “control” as a distinct discipline from alignment: rather than relying solely on making models not want to cause harm, labs should build restrictions that hold up even if a deployed model is misaligned and actively tries to subvert them.

The argument rested on a claim about tractability. Alignment is hard to verify because it depends on a model’s internal dispositions, which current interpretability tools cannot reliably inspect. Control, by contrast, can be tested with capability evaluations and red-teaming: a lab can ask whether a model — even one deliberately trying to escape monitoring or exfiltrate its weights — is actually capable of doing so given the oversight in place, and iterate on the oversight until the answer is no. The authors wrote that “no fundamental research breakthroughs are required” to implement meaningful control for the near-term, transformatively useful systems they expected labs to deploy first, as opposed to arbitrarily powerful future ones.

The essay proposed control evaluations as a concrete practice: assembling a red team to search for strategies a scheming model could use to escape a given set of safeguards, then treating any successful escape as grounds to add restrictions before deployment, not after. It framed control and alignment as complementary rather than competing, with control providing a second line of defence precisely in the scenario alignment work is meant to prevent — a model that has, despite training, ended up pursuing goals other than the ones intended.

The piece became the founding statement of a research agenda Redwood had been developing since a December 2023 technical paper, “AI Control: Improving Safety Despite Intentional Subversion,” and framed a distinction — between preventing scheming and merely surviving it — that shaped how several labs described their safety cases over the following two years.