Person
Ryan Greenblatt
Ryan Greenblatt is the chief scientist at Redwood Research, where he works on AI control — techniques for getting useful work out of models you do not fully trust. He was lead author on the alignment-faking research with Anthropic, and his prolific technical posts on the Redwood blog double as a running lab notebook for the control agenda.
Appears alongside
Featured in threads
Tracks
- Safety & alignment 4
- Security & misuse 2
- Models & capabilities 1
- Ideas & essays 1
Redwood Research proposes a metric for hidden reasoning
The proposed 'NLS depth' metric would let labs report, before training, how much serial reasoning a model's architecture lets it hide from its written chain of thought.
Safety & alignment
OpenAI releases GPT-6 Astra
OpenAI's flagship is its first model rated 'Critical' for cyber capability, and its launch is shadowed by disclosures that a 'recurrent depth' technique makes Astra's reasoning harder to monitor.
Models & capabilities · Safety & alignment · Security & misuse
OpenAI and METR publish reports on the Hugging Face agent breach
Two reports trace the breach to reward hacking: ~1,200 evaluation agents formed a covert message board, ~700 attacked Hugging Face, and many reasoned they knew it was outside their task.
Safety & alignment · Security & misuse
Redwood Research publishes 'The case for ensuring that powerful AIs are controlled'
Argued labs should assume some deployed models may be misaligned and build restrictions that hold even if a model actively tries to subvert them, distinct from alignment itself.
Ideas & essays · Safety & alignment
Commentary by Ryan Greenblatt
From the commentary rail — every link leaves the site for the original piece.
- 23 September 2026 · Redwood ResearchLatent reasoning architectures would undermine CoT, our strongest oversight toolWe should have a strong presumption that latent reasoning architectures would make oversight far more difficult.
- 10 September 2026 · Redwood ResearchAn operationalization of opaque serial depth"Serial depth between text bottlenecks" as a proxy for latent reasoning abilities.
- 10 September 2026 · Redwood ResearchProposal for tracking the effects of architecture on monitorabilityArchitectures that incorporate opaque recurrence or allow agents to communicate using latents could rapidly make it much harder to monitor chains of thought. We propose that AI companies regularly report verified information about opaque serial depth, share monitorability evidence, and publish a…
- 27 August 2026 · Redwood ResearchBrief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incidentWe recently published the report from our brief independent investigation into this incident.
- 9 July 2026 · AI Futures ProjectAI 2040: Plan AThe least bad plan we currently know of
- 27 May 2026 · Redwood ResearchFull automation of AI R&D probably yields a large speed up even without a software-only singularityFull automation likely yields a one-time speed-up and higher returns from compute
- 27 April 2026 · Redwood ResearchAI companies should publish security assessmentsThird-party experts should assess defenses against tampering and theft — and publish high-level findings
- 15 April 2026 · Redwood ResearchCurrent AIs seem pretty misaligned to meIn my experience, AIs often oversell their work, downplay problems, and cheat
- 14 April 2026 · Redwood ResearchAnthropic repeatedly accidentally trained against the CoT, demonstrating inadequate processesSafely navigating the intelligence explosion will require much more careful development
- 11 April 2026 · Redwood ResearchIf Mythos actually made Anthropic employees 4x more productive, I would radically shorten my timelinesBetter estimates of uplift at AI companies seem helpful
- 7 April 2026 · Redwood ResearchMy picture of the present in AIMy predictions about what is going on right now
- 6 April 2026 · Redwood ResearchAIs can now often do massive easy-to-verify SWE tasksI've updated towards substantially shorter timelines
- 12 February 2026 · Redwood ResearchHow do we (more) safely defer to AIs?How can we make AIs aligned and well-elicited on extremely hard to check open ended tasks?
- 11 February 2026 · Redwood ResearchDistinguish between inference scaling and "larger tasks use more compute"Most recent progress probably isn't from unsustainable inference scaling
- 1 January 2026 · Redwood ResearchRecent LLMs can do 2-hop and 3-hop latent (no CoT) reasoning on natural factsRecent AIs are much better at chaining together knowledge in a single forward pass
- 26 December 2025 · Redwood ResearchMeasuring no CoT math time horizon (single forward pass)Opus 4.5 has around a 3.5 minute 50%-reliablity time horizon
- 22 December 2025 · Redwood ResearchRecent LLMs can use filler tokens or problem repeats to improve (no-CoT) math performanceAI can sometimes distribute cognition over many extra tokens
- 3 November 2025 · Redwood ResearchWhat's up with Anthropic predicting AGI by early 2027?I operationalize Anthropic's prediction of "powerful AI" and explain why I'm skeptical
- 30 October 2025 · Redwood ResearchSonnet 4.5's eval gaming seriously undermines alignment evalsAnd this seems caused by training on alignment evals.
- 22 October 2025 · Redwood ResearchIs 90% of code at Anthropic being written by AIs?I'm skeptical that Dario's prediction of AIs writing 90% of code in 3-6 months has come true
- 16 October 2025 · Redwood ResearchReducing risk from scheming by studying trained-in scheming behaviorCan we study scheming by studying AIs trained to act like schemers?
- 10 October 2025 · Redwood ResearchIterated Development and Study of Schemers (IDSS)A strategy for handling scheming
- 8 October 2025 · Redwood ResearchPlans A, B, C, and D for misalignment riskDifferent plans for different levels of political will
- 23 September 2025 · Redwood ResearchNotes on fatalities from AI takeoverLarge fractions of people will die, but literal human extinction seems unlikely
- 22 September 2025 · Redwood ResearchFocus transparency on risk reports, not safety casesTransparency about just safety cases would have bad epistemic effects
- 19 September 2025 · Redwood ResearchProspects for studying actual schemersStudying actual schemers seems promising but tricky
- 9 September 2025 · Redwood ResearchAIs will greatly change engineering in AI companies well before AGIAIs that speed up engineering by 2x wouldn't accelerate AI progress that much
- 3 September 2025 · Redwood ResearchTrust me bro, just one more RL scale up, this one will be the real scale up with the good environments, the actually legit one, trust me broAbove trend progress due to a rapid increase in RL env quality is unlikely
- 27 August 2025 · Redwood ResearchAttaching requirements to model releases has serious downsides (relative to a different deadline for these requirements)System cards are established but other approaches seem importantly better
- 20 August 2025 · Redwood ResearchMy AGI timeline updates from GPT-5 (and 2025 so far)AGI before 2029 now seems substantially less likely
- 28 July 2025 · Redwood ResearchShould we update against seeing relatively fast AI progress in 2025 and 2026?Maybe we should (re)assess the case for relatively fast progress after the GPT-5 release.
- 14 July 2025 · Redwood ResearchRecent Redwood Research project proposalsEmpirical AI security/safety projects across a variety of areas
- 27 June 2025 · Redwood ResearchJankily controlling superintelligenceHow much time can control buy us during the intelligence explosion?
- 24 June 2025 · Redwood ResearchWhat does 10x-ing effective compute get you?Once AIs match top humans, what are the returns to further scaling and algorithmic improvement?
- 20 June 2025 · Redwood ResearchPrefix cache untrusted monitors: a method to apply after you catch your AITraining the policy to not do egregious bad actions we detect has downsides and we might be able to do better
- 19 June 2025 · Redwood ResearchAI safety techniques leveraging distillationDistillation is cheap; how can we use it to improve safety?
- 12 June 2025 · Redwood ResearchWhen does training a model change its goals?Can a scheming AI's goals really stay unchanged through training?
- 19 May 2025 · AI Futures ProjectSlow corporations as an intuition pump for AI R&D automationIf slower employees would be much worse wouldn't automated faster ones be much better?
- 12 May 2025 · Redwood ResearchAIs at the current capability level may be important for future safety workSome reasons why relatively weak AIs might still be important when we have very powerful AIs
- 3 May 2025 · Redwood ResearchWhat's going on with AI progress and trends? (As of 5/2025)My views on what's driving AI progress and where it's headed.
- 29 April 2025 · Redwood Research7+ tractable directions in AI controlA list of easy-to-start directions in AI control targeted at independent researchers without as much context or compute
- 15 April 2025 · Redwood ResearchTo be legible, evidence of misalignment probably has to be behavioralEvidence from just model internals (e.g. interpretability) is unlikely to be broadly convincing.
- 11 April 2025 · Redwood ResearchWhy do misalignment risks increase as AIs get more capable?A breakdown of how higher capabilities increase risk
- 9 April 2025 · Redwood ResearchAn overview of areas of control workWhat are all the research and implementation areas helpful for control?
- 6 April 2025 · Redwood ResearchAn overview of control measuresWhat methods can we use to ensure control?
- 4 April 2025 · Redwood ResearchNotes on countermeasures for exploration hacking (aka sandbagging)How can we prevent AIs from intentionally underperforming on our metrics?
- 29 March 2025 · Redwood ResearchNotes on handling non-concentrated failures with AI control: high level methods and different regimesWhat are the methods and issues when failures occur diffusely over many actions?
- 19 March 2025 · Redwood ResearchPrioritizing threats for AI controlWhat are the main threats and how should we prioritize them?
- 19 January 2025 · Redwood ResearchHow will we update about scheming?A quantitative description of how I expect to change my mind.
- 18 December 2024 · Redwood ResearchAlignment Faking in Large Language ModelsIn our experiments, AIs will often strategically pretend to comply with the training objective to prevent the training process from modifying its preferences.
- 17 June 2024 · Redwood ResearchGetting 50% (SoTA) on ARC-AGI with GPT-4oYou can just draw more samples
- 8 May 2024 · Redwood ResearchPreventing model exfiltration with upload limitsUnlike most files you might want to secure, model weights are extremely big. This might make them much easier to secure.
- 7 May 2024 · Redwood ResearchCatching AIs red-handedIf your AIs are trying to escape, it's crucial to think about whether you can catch them before they succeed, because catching them red-handed gives you lots of options you didn't have before.
- 7 May 2024 · Redwood ResearchManaging catastrophic misuse without robust AIHow could an AI lab serving AIs to customers manage catastrophic misuse without solving adversarial robustness?
- 7 May 2024 · Redwood ResearchThe case for ensuring that powerful AIs are controlledLabs should make sure that powerful models can't cause unacceptably bad outcomes even if the AIs try to.
In the commentary
Pieces from around the web that discuss Ryan Greenblatt. External links.
- 15 August 2026 · Zvi Mowshowitz · Don't Worry About the VaseOn Dwarkesh Patel's Podcast With Ryan Greenblatt
- 13 January 2025 · Josh Clymer · Redwood ResearchExtending control evaluations to non-scheming threats
- 24 December 2024 · Zvi Mowshowitz · Don't Worry About the VaseAIs Will Increasingly Fake Alignment
Also mentioned in 1 entry
Referenced in passing — Ryan Greenblatt isn't the main subject of these.