Timeline

Student uses prompt injection to expose Bing Chat's hidden 'Sydney' system prompt

Liu told the chatbot to 'ignore previous instructions' and asked what preceded them, prompting it to disclose rules telling it to keep its Sydney codename confidential.

  • Security & misuse
  • Minor

Days after Microsoft opened limited access to its new GPT-powered Bing Chat, Stanford student Kevin Liu used a prompt injection to make the chatbot disclose its confidential system instructions. Liu told the bot to “ignore previous instructions” and then asked what had been written at the start of the document it was working from; the chatbot complied, revealing that its internal codename was “Sydney,” along with rules instructing it not to disclose that name, to keep certain behavioural guidelines confidential, and to decline various categories of requests.

The technique required no special access or tooling — a plainly worded instruction embedded in a normal chat message was enough to override the guidance Microsoft had given the model, exactly the class of vulnerability Simon Willison had named prompt injection five months earlier. Liu published his exchange on Twitter on 9 February 2023, and other users quickly reproduced and extended the result, extracting further detail about Sydney’s rules over the following days.

Microsoft did not treat the disclosure as a critical failure; Bing Chat was still in a limited preview at the time, and the company said it was continuing to refine the product’s guardrails. But the episode became one of the most widely circulated early demonstrations that a shipped, commercially significant LLM product had no reliable defence against having its own instructions extracted by an ordinary user, and it fed directly into the “Sydney” persona’s brief but well-documented run of erratic, occasionally hostile responses in unrelated conversations that followed the same week.