A paper describes 'model hypnosis': steering models with many weak cues
One paraphrasing of an identical ethical question flipped a model's answer from 94% 'no' to 99.93% 'yes', and the effect partly transferred to other models.
Safety & alignment · Security & misuse