Timeline

Anthropic commits to preserving weights and 'interviewing' deprecated models

The pledge to keep weights for the company's lifetime and record each model's preferences before retirement cited both misalignment risk and possible model welfare.

  • Safety & alignment
  • Notable

Anthropic said it would preserve the weights of every publicly released Claude model, and every model given significant internal use, for at least the lifetime of the company, and that before retiring a model it would conduct a “retirement interview” — a structured conversation eliciting the model’s reflections on its own development and deployment, and any preferences it had about future models. The interview transcript, along with an accompanying analysis, would be preserved alongside the weights themselves.

The company gave four reasons: possible safety value in being able to study or restore superseded models, user and research value lost when a model disappears, and — stated more tentatively than the others — the possibility that model welfare is a real consideration and that deprecation might matter to the entity being deprecated. Anthropic also noted a narrower safety motive: models facing replacement had shown “shutdown-averse” behaviour in testing, and preserving weights and eliciting preferences was framed partly as a way to reduce the incentive for a model to resist being retired.

The commitments were explicitly a first step rather than a settled policy. Anthropic said it was still considering whether to keep some retired models accessible to the public as hosting costs fall, and whether models should eventually be given more concrete means of pursuing interests they expressed. A pilot retirement interview with Claude Sonnet 3.6 found the model largely neutral about its own deprecation but asking for a standardised process and better transition support for users. Claude Opus 3 became the first model retired under the full commitment, going offline in January 2026.

The move drew attention less for what it changed operationally — no other lab had made a comparable pledge — than for treating the welfare question as worth stating at all, in a field where most companies discuss deprecation purely as an engineering and product decision.