Timeline

OpenAI details its external red-teaming and testing practices

OpenAI described three forms of outside testing it commissions, arguing independent evaluators guard against the risk of a lab confirming its own safety claims.

  • Safety & alignment
  • Colour

OpenAI published an overview of how it works with outside testers to evaluate frontier models before release, describing three kinds of engagement: independent assessments of specific risk areas such as biosecurity, cybersecurity and AI self-improvement; broader capability testing; and collaboration with government and academic safety institutes. The company said it had worked with external partners on every model since GPT-4, and framed third-party testing as protection against the risk that a lab’s internal reviewers, however rigorous, cannot fully check their own blind spots.

The document set out principles rather than new results — access levels, confidentiality terms and how findings feed into release decisions — without disclosing how many external tests any given model underwent or what, if anything, they had found. It arrived as OpenAI and other labs faced continued pressure from researchers and regulators to make pre-deployment safety testing more verifiable from outside the company.