Person
Chris Olah
Chris Olah is a co-founder of the AI lab Anthropic and one of the leading figures in mechanistic interpretability — the effort to understand what is actually happening inside a neural network rather than treating it as a black box. He leads Anthropic's interpretability team, whose work on "features" and sparse autoencoders (in papers such as "Towards Monosemanticity" and "Scaling Monosemanticity") aims to identify the human-readable concepts a model represents internally. The argument for this research is that being able to inspect a system's inner workings is a prerequisite for keeping powerful models safe and controllable. He is known for unusually clear public explanations of technical ideas, and has argued for the value of critics who sit outside the commercial pressures the labs face.
Featured in threads
Tracks
- Safety & alignment 4
- Culture & impact 1
- Ideas & essays 1
Anthropic co-founder Chris Olah meets Pope Leo on AI encyclical
Olah told the Vatican that AI companies operate under commercial and geopolitical pressures that can conflict with safety, and argued for outside critics 'who care about things going well.'
Culture & impact
Anthropic maps millions of concepts inside a production model
Sparse autoencoders extracted human-interpretable features from a deployed model — and turning one up produced Golden Gate Claude.
Safety & alignment · Ideas & essays
Anthropic publishes 'In-Context Learning and Induction Heads'
Anthropic's interpretability team argued a single attention mechanism, found across model sizes, does most of the work behind a model's ability to learn from its prompt.
Safety & alignment
Anthropic publishes 'A Mathematical Framework for Transformer Circuits'
Studying deliberately simplified transformers with no more than two layers, the team found 'induction heads' — a mechanism later argued to explain much of in-context learning.
Safety & alignment
OpenAI publishes multimodal neurons research
A single neuron fired for photos of spiders, drawings of Spider-Man and the rendered word 'spider' alike — and pasting a mislabelled sticker on an object could fool the model.
Safety & alignment
Also mentioned in 2 entries
Referenced in passing — Chris Olah isn't the main subject of these.