Timeline

A paper describes 'model hypnosis': steering models with many weak cues

One paraphrasing of an identical ethical question flipped a model's answer from 94% 'no' to 99.93% 'yes', and the effect partly transferred to other models.

  • Safety & alignment
  • Security & misuse
  • Minor

Enric Boix-Adsera and Benedict Tessler, of the University of Pennsylvania, described what they called “model hypnosis”: the finding that individually weak, seemingly irrelevant cues in a prompt — the choice of animals mentioned in unrelated context, a meaning-preserving paraphrase, scattered typos, or the values in a JSON metadata block — can be combined systematically to strongly steer a language model’s output on an unrelated question, even though no single cue functions as an instruction or provides relevant evidence.

The paper’s headline example held the substance of an ethical question fixed and varied only its phrasing: one version produced “no” with 94% probability, an alternative phrasing of the identical question produced “yes” with 99.93% probability. The authors reported this pattern held across four cue families and across model families and scales, including in frontier reasoning models, and that hypnotic prompts optimised against one model often retained a directional effect when tested against a different one — evidence, they argued, that models share some of the same underlying response biases rather than each having idiosyncratic quirks.

The authors framed the result as a safety and interpretability problem rather than a demonstrated attack: a prompt built from many additive weak signals is harder to flag than a single adversarial trigger phrase, since no individual component looks suspicious on inspection. They described it as opening “new challenges and avenues for AI safety” and called it “a major hurdle for AI interpretability.” The paper’s abstract page does not detail a proposed defence.