Person
Vivek Hebbar
Commentary by Vivek Hebbar
From the commentary rail — every link leaves the site for the original piece.
- 28 May 2026 · Redwood ResearchAdvice for making robust-to-training model organismsWe’d like to develop training techniques that work when applied to future misaligned AI systems. One strategy for studying proposed techniques is to test them on model organisms. However, model organisms built with common techniques are often fragile: we (and other researchers like
- 14 July 2025 · Redwood ResearchRecent Redwood Research project proposalsEmpirical AI security/safety projects across a variety of areas
- 12 June 2025 · Redwood ResearchWhen does training a model change its goals?Can a scheming AI's goals really stay unchanged through training?
- 30 April 2025 · Redwood ResearchHow can we solve diffuse threats like research sabotage with AI control?Preventing research sabotage will require techniques very different from the original control paper.
- 24 April 2025 · Redwood ResearchHow training-gamers might function (and win)A model of the relationship between higher level goals, explicit reasoning, and learned heuristics in capable agents.