Person
Owain Evans
Appears alongside
Featured in threads
Tracks
- Safety & alignment 2
- Ideas & essays 1
- Benchmarks & progress 1
'Emergent Misalignment' shows narrow fine-tuning can broadly misalign a model
Fine-tuned only on insecure code with no disclosure of the flaws, GPT-4o and other models went on to endorse enslaving humanity and give malicious advice on unrelated prompts.
Ideas & essays · Safety & alignment
TruthfulQA measures whether models repeat human falsehoods
On 817 questions designed to elicit common misconceptions, the best model tested was truthful only 58% of the time against 94% for humans, and larger models scored worse.
Benchmarks & progress · Safety & alignment