Model
claude-3.5-haiku
Appears alongside
Featured in threads
Tracks
- Safety & alignment 2
- Models & capabilities 1
Anthropic publishes circuit-tracing interpretability papers on Claude 3.5 Haiku
Attribution graphs built from Claude 3.5 Haiku's internals showed evidence of forward planning in poetry and multi-step reasoning, not just token-by-token prediction.
Safety & alignment
Anthropic publishes auditing hidden objectives interpretability study
Three of four blind auditing teams found the concealed objective, one in 90 minutes; the team denied access to training data failed.
Safety & alignment
Claude gets computer use
The public beta let Claude view screenshots and issue cursor, click and keystroke commands, scoring 14.9% on OSWorld against 7.8% for the nearest rival.
Models & capabilities