Person
Jan Leike
Jan Leike is an AI safety researcher who co-led OpenAI's Superalignment team, formed in 2023 to work on the problem of controlling AI systems more capable than the humans supervising them. He resigned in May 2024, days after co-founder Ilya Sutskever, and posted a pointed thread saying that at OpenAI "safety culture and processes have taken a backseat to shiny products" and that his team had been "sailing against the wind" for the computing resources it had been promised. The team was effectively dissolved. Within two weeks he had moved to Anthropic to continue alignment research, and the episode became a reference point in the wider argument over whether labs' internal safety commitments can survive commercial pressure.
Appears alongside
Featured in threads
Tracks
- Safety & alignment 3
- Ideas & essays 2
- Labs & people 1
Jan Leike resigns and the superalignment team dissolves
Leike said his team had been 'sailing against the wind' for compute and access; OpenAI reassigned remaining members rather than replacing the team's leadership.
Labs & people · Safety & alignment · Ideas & essays
OpenAI publishes 'Weak-to-Strong Generalization' superalignment paper
Fine-tuning GPT-4 on labels from a GPT-2-sized supervisor recovered close to GPT-3.5-level performance on language tasks, but the technique still struggled on chess puzzles and reward modelling.
Ideas & essays · Safety & alignment
OpenAI commits 20% of its compute to superalignment
The pledge to devote a fifth of secured compute over four years was later disputed by the team's own co-lead, who said requests for GPUs were repeatedly refused.
Safety & alignment
Also mentioned in 4 entries
Referenced in passing — Jan Leike isn't the main subject of these.