Person
Christopher Olah
Appears alongside
Featured in threads
Tracks
- Safety & alignment 2
Anthropic publishes 'Towards Monosemanticity'
Sparse autoencoders decomposed a single 512-neuron layer into more than 4,000 human-interpretable features, far more than the raw neurons showed.
Safety & alignment
Anthropic publishes 'Toy Models of Superposition'
Elhage, Olah and colleagues showed small networks represent more features than they have neurons by packing them into overlapping directions, complicating efforts to read a model's internals.
Safety & alignment