Person
Buck Shlegeris
Buck Shlegeris is the chief executive of Redwood Research, a Berkeley nonprofit founded in 2021 that studies technical approaches to controlling powerful AI systems. He co-founded the organisation as its chief technology officer before becoming CEO in 2024, and previously worked at the Machine Intelligence Research Institute.
Appears alongside
Featured in threads
Tracks
- Safety & alignment 2
- Models & capabilities 1
- Security & misuse 1
- Ideas & essays 1
OpenAI releases GPT-6 Astra
OpenAI's flagship is its first model rated 'Critical' for cyber capability, and its launch is shadowed by disclosures that a 'recurrent depth' technique makes Astra's reasoning harder to monitor.
Models & capabilities · Safety & alignment · Security & misuse
Redwood Research publishes 'The case for ensuring that powerful AIs are controlled'
Argued labs should assume some deployed models may be misaligned and build restrictions that hold even if a model actively tries to subvert them, distinct from alignment itself.
Ideas & essays · Safety & alignment
Commentary by Buck Shlegeris
From the commentary rail — every link leaves the site for the original piece.
- 8 June 2026 · Redwood ResearchEfficient tradeoffs and the safety-usefulness tradeoff modelWhen is "increasing safety budget" a useful concept?
- 11 May 2026 · Redwood ResearchHow useful is the information you get from working inside an AI company?My median guess: it's as good as a crystal ball that sees 2.5 months into the future.
- 7 May 2026 · Redwood ResearchA review of “Investigating the consequences of accidentally grading CoT during RL”Last week, OpenAI staff shared an early draft of Investigating the consequences of accidentally grading CoT during RL with Redwood Research staff.
- 4 December 2025 · Redwood ResearchThe behavioral selection model for predicting AI motivationsThe basic arguments about AI motivations in one causal graph
- 9 October 2025 · Redwood ResearchThe Thinking Machines Tinker API is good news for AI control and securityIt's a promising design for reducing model access inside AI companies.
- 9 August 2025 · Redwood ResearchFour places where you can put LLM monitoringTo wit: LLM APIs, agent scaffolds, code review, and detection-and-response systems
- 18 July 2025 · Redwood ResearchWhy it's hard to make settings for high-stakes control researchIt's like making challenging evals, but more constrained
- 14 July 2025 · Redwood ResearchRecent Redwood Research project proposalsEmpirical AI security/safety projects across a variety of areas
- 9 July 2025 · Redwood ResearchWhat's worse, spies or schemers?And what if you have both at once?
- 5 July 2025 · Redwood ResearchHow much novel security-critical infrastructure do you need during the singularity?And what does this mean for AI control?
- 2 July 2025 · Redwood ResearchThere are two fundamentally different constraints on schemers"They need to act aligned" often isn't precise enough
- 23 June 2025 · Redwood ResearchComparing risk from internally-deployed AI to insider and outsider threats from humansAnd why I think insider threat from AI combines the hard parts of both problems.
- 20 June 2025 · Redwood ResearchMaking deals with early schemers...could help us to prevent takeover attempts from more dangerous misaligned AIs created later.
- 8 May 2025 · Redwood ResearchMisalignment and Strategic Underperformance: An Analysis of Sandbagging and Exploration HackingA new analysis of the risk of AIs intentionally performing poorly.
- 18 April 2025 · Redwood ResearchHandling schemers if shutdown is not an optionWhat if getting strong evidence of scheming isn't the end of your scheming problems, but merely the middle?
- 16 April 2025 · Redwood ResearchCtrl-Z: Controlling AI Agents via ResamplingA new paper on AI control for agents.
- 28 January 2025 · Redwood ResearchTen people on the insideA scary scenario that's worth planning for
- 17 January 2025 · Redwood ResearchThoughts on the conservative assumptions in AI controlWhy are we so friendly to the red team?
- 20 December 2024 · Redwood ResearchMeasuring whether AIs can statelessly strategize to subvert security measuresThe complement to control evaluations
- 18 December 2024 · Redwood ResearchAlignment Faking in Large Language ModelsIn our experiments, AIs will often strategically pretend to comply with the training objective to prevent the training process from modifying its preferences.
- 18 November 2024 · Redwood ResearchWhy imperfect adversarial robustness doesn't doom AI controlThere are crucial disanalogies between preventing jailbreaks and preventing misalignment-induced catastrophes.
- 15 November 2024 · Redwood ResearchWin/continue/lose scenarios and execute/replace/audit protocolsIn this post, I’ll make a technical point that comes up when thinking about risks from scheming AIs from a control perspective.
- 10 October 2024 · Redwood ResearchBehavioral red-teaming is unlikely to produce clear, strong evidence that models aren't schemingOne strategy for mitigating risk from schemers (that is, egregiously misaligned models that intentionally try to subvert your safety measures) is behavioral red-teaming (BRT).
- 26 September 2024 · Redwood ResearchA basic systems architecture for AI agents that do autonomous researchAnd diagrams describing how threat scenarios involving misaligned AI involve compromising the system in different places.
- 25 September 2024 · Redwood ResearchHow to prevent collusion when using untrusted models to monitor each otherSuppose you’ve trained a really clever AI model, and you’re planning to deploy it in an agent scaffold that allows it to run code or take other actions.
- 26 August 2024 · Redwood ResearchWould catching your AIs trying to escape convince AI developers to slow down or undeploy?I'm not so sure.
- 13 August 2024 · Redwood ResearchFields that I reference when thinking about AI takeover preventionIs AI takeover like a nuclear meltdown? A coup? A plane crash?
- 10 June 2024 · Redwood ResearchAccess to powerful AI might make computer security radically easierAI might be really helpful for reducing security risk.
- 3 June 2024 · Redwood ResearchAI catastrophes and rogue deploymentsIt’s interesting to classify possible AI catastrophes based on whether or not they involve a "rogue deployment".
- 7 May 2024 · Redwood ResearchCatching AIs red-handedIf your AIs are trying to escape, it's crucial to think about whether you can catch them before they succeed, because catching them red-handed gives you lots of options you didn't have before.
- 7 May 2024 · Redwood ResearchManaging catastrophic misuse without robust AIHow could an AI lab serving AIs to customers manage catastrophic misuse without solving adversarial robustness?
- 7 May 2024 · Redwood ResearchThe case for ensuring that powerful AIs are controlledLabs should make sure that powerful models can't cause unacceptably bad outcomes even if the AIs try to.
- 7 May 2024 · Redwood ResearchUntrusted smart models and trusted dumb modelsThe easiest way to be sure an AI isn't scheming against you is to note that it's too dumb to pull that off. What happens if that's the only way we have to rule out scheming?
In the commentary
Pieces from around the web that discuss Buck Shlegeris. External links.