Person
Sebastian Prasanna
Commentary by Sebastian Prasanna
From the commentary rail — every link leaves the site for the original piece.
- 23 September 2026 · Redwood ResearchAstra is much better at reasoning with filler tokens than previous modelsUnlike other models we tested, GPT-6 Astra performs modestly better on general benchmarks with filler tokens, and significantly better on serial depth-heavy tasks.
- 18 June 2026 · Redwood ResearchThe distillation double bind: Distilling misaligned models either transfers misalignment or it doesn'tIf it transfers misalignment, we might get a misaligned model that’s easier to incriminate. If it doesn’t, we might get a capable benign replacement model.
- 28 May 2026 · Redwood ResearchAdvice for making robust-to-training model organismsWe’d like to develop training techniques that work when applied to future misaligned AI systems. One strategy for studying proposed techniques is to test them on model organisms. However, model organisms built with common techniques are often fragile: we (and other researchers like
- 18 May 2026 · Redwood ResearchIncriminating misaligned AI models via distillationSuppose we have a dangerous misaligned AI that can fool alignment audits, and distill it into a student model.