Person

Chris Olah

5 entries · March 2021 – May 2026

Chris Olah is a co-founder of the AI lab Anthropic and one of the leading figures in mechanistic interpretability — the effort to understand what is actually happening inside a neural network rather than treating it as a black box. He leads Anthropic's interpretability team, whose work on "features" and sparse autoencoders (in papers such as "Towards Monosemanticity" and "Scaling Monosemanticity") aims to identify the human-readable concepts a model represents internally. The argument for this research is that being able to inspect a system's inner workings is a prerequisite for keeping powerful models safe and controllable. He is known for unusually clear public explanations of technical ideas, and has argued for the value of critics who sit outside the commercial pressures the labs face.

Featured in threads

Tracks

  • Safety & alignment 4
  • Culture & impact 1
  • Ideas & essays 1

Also mentioned in 2 entries

Referenced in passing — Chris Olah isn't the main subject of these.