Person
Kyle Fish
Appears alongside
Featured in threads
Tracks
- Safety & alignment 2
Anthropic maps the 'Assistant Axis' persona vector across open models
An intervention called activation capping, which constrains a model's activations to normal range, cut harmful persona-drift responses by roughly half in testing.
Safety & alignment
Anthropic hires its first AI welfare researcher
Fish, a co-author of the 'Taking AI Welfare Seriously' report, joined Anthropic's alignment science team; the company's public statement on model welfare followed roughly six weeks later.
Safety & alignment