OpenAI publishes Operator system card
External red-teamers targeted prompt injection specifically; OpenAI reported raising its injection-detection recall from 79% to 99% after one testing round.
- Safety & alignment
- Minor
OpenAI published the system card accompanying Operator’s launch, its safety evaluation of the Computer-Using Agent model. The card identified prompt injection — adversarial instructions embedded in a webpage’s visible or hidden content — as the most significant new risk category the product introduced, since a browsing agent could encounter and act on hostile text anywhere on the open web, not only in what a user typed.
OpenAI said external red-teamers tested the model in two phases, first without safety mitigations and then against the mitigated version, probing prompt injection and attempts to override the system’s instructions using spoofed emails and malicious pages in test environments. The company reported building a dedicated injection classifier and, after one round of red-teaming surfaced new attack patterns, raising its detection recall from 79% to 99% within a day. It also described training the model to recognise and refuse manipulative page content, and requiring explicit user confirmation before financial transactions, sending emails, or deleting data.
Operator was rated low risk on OpenAI’s biological and autonomous-replication evaluations — under 10% task success on autonomy benchmarks, with biorisk tooling limited largely by the model’s poor optical character recognition on complex strings such as DNA sequences — and refused around 97% of tasks flagged as harmful in agent-specific test scenarios. The card became a reference point for the computer-use system cards other labs published as browsing agents became more common through 2025.