Timeline

Anthropic publishes its first alignment paper

'A General Language Assistant as a Laboratory for Alignment' introduced the helpful-honest-harmless framing and found preference modelling scales better than imitation.

  • Safety & alignment
  • Ideas & essays
  • Minor

Anthropic, founded six months earlier by researchers who had left OpenAI, published its first alignment research paper, “A General Language Assistant as a Laboratory for Alignment.” The paper set out simple baseline techniques for building a text-based assistant aligned with human values, framed around a target the company would keep using for years afterward: that a model should be helpful, honest and harmless — the “HHH” criteria.

The paper’s central empirical result concerned how best to incorporate human judgement into training. Comparing imitation learning, binary discrimination and ranked preference modelling — where a model learns to rank several candidate responses rather than simply copy a preferred one — the authors found that preference modelling substantially outperformed imitation learning, and that this advantage grew, rather than shrank, as models got larger. They also reported that a preference-model pre-training phase improved the sample efficiency of later fine-tuning on human feedback, and that these interventions did not degrade the model’s general capabilities.

The paper carried little of the capability or product news that would characterise the industry’s later output, but it stated the research direction — preference-based fine-tuning, scaled with model size — that Anthropic’s subsequent work on reinforcement learning from human feedback and Constitutional AI would build on directly. The HHH framing itself outlived the paper’s specific experiments, recurring in Anthropic’s public description of its models for years afterward.