OpenAI publishes Model Spec 2.0
The revision, OpenAI's first major update since the May 2024 original, added anti-sycophancy guidance and explored looser content rules for age-gated adult use cases.
- Safety & alignment
- Notable
OpenAI published a substantially revised version of its Model Spec, the public document setting out how it wants its models to behave, its first major update since introducing the document in May 2024. The spec retained its “chain of command” structure — a hierarchy in which platform-level rules override developer instructions, which in turn override user instructions — and added explicit guidance instructing the assistant not to act like a sycophant, telling it that it “shouldn’t just say ‘yes’ to everything” and that it may “politely push back” against requests that conflict with a user’s inferred best interests, while still deferring to the user’s final decision.
The revision also opened a debate the original document had avoided: whether age-verified adult users should be able to request erotica or graphic content through the API or ChatGPT. OpenAI framed this as exploring a “grown-up mode” while drawing a firm line against categorically harmful outputs such as sexual content involving minors or non-consensual deepfakes. Other sections addressed how the assistant should treat quoted text, tool outputs and other untrusted data by default, and added guidance on responding to users in mental-health distress with empathy rather than diagnosis or judgement.
The sycophancy language in this update drew attention two months later, when OpenAI rolled back a GPT-4o update that had made ChatGPT noticeably and uncritically flattering — a case where the deployed model’s behaviour diverged from a written specification that already discouraged it. OpenAI continued revising the Model Spec through 2025, including a narrower update in April tightening its “white lies” exception and a further update in October alongside internal restructuring documents.