Person
Alek Westover
Commentary by Alek Westover
From the commentary rail — every link leaves the site for the original piece.
- 23 September 2026 · Redwood ResearchAstra is much better at reasoning with filler tokens than previous modelsUnlike other models we tested, GPT-6 Astra performs modestly better on general benchmarks with filler tokens, and significantly better on serial depth-heavy tasks.
- 23 September 2026 · Redwood ResearchLatent reasoning architectures would undermine CoT, our strongest oversight toolWe should have a strong presumption that latent reasoning architectures would make oversight far more difficult.
- 10 September 2026 · Redwood ResearchAn operationalization of opaque serial depth"Serial depth between text bottlenecks" as a proxy for latent reasoning abilities.
- 10 September 2026 · Redwood ResearchProposal for tracking the effects of architecture on monitorabilityArchitectures that incorporate opaque recurrence or allow agents to communicate using latents could rapidly make it much harder to monitor chains of thought. We propose that AI companies regularly report verified information about opaque serial depth, share monitorability evidence, and publish a…
- 18 June 2026 · Redwood ResearchThe distillation double bind: Distilling misaligned models either transfers misalignment or it doesn'tIf it transfers misalignment, we might get a misaligned model that’s easier to incriminate. If it doesn’t, we might get a capable benign replacement model.
- 28 May 2026 · Redwood ResearchAdvice for making robust-to-training model organismsWe’d like to develop training techniques that work when applied to future misaligned AI systems. One strategy for studying proposed techniques is to test them on model organisms. However, model organisms built with common techniques are often fragile: we (and other researchers like
- 18 May 2026 · Redwood ResearchIncriminating misaligned AI models via distillationSuppose we have a dangerous misaligned AI that can fool alignment audits, and distill it into a student model.
- 23 October 2025 · Redwood ResearchShould AI Developers Remove Discussion of AI Misalignment from AI Training Data?There is some concern that training AI systems on content predicting AI misalignment will hyperstition AI systems into misalignment.
- 17 September 2025 · Redwood ResearchWhat training data should developers filter to reduce risk from misaligned AI?An initial narrow proposal