Timeline

Nvidia and Anthropic launch an open platform to contain AI agents

Nvidia said over 100 organisations, including Anthropic and Microsoft, were using OpenShell software and a hardware watchdog called Sentry to enforce agent boundaries from outside the model.

  • Security & misuse
  • Compute & infrastructure
  • Notable

Nvidia and Anthropic launched the Open Agent Safety Platform, a set of tools intended to stop AI agents crossing security boundaries regardless of what the underlying model itself decides to do. Nvidia said over 100 organisations were working with the platform at launch, including Anthropic, Microsoft, Hugging Face, Cisco, CrowdStrike, Palantir and Salesforce.

The platform has two parts. OpenShell is open-source runtime software, licensed under Apache 2.0, that runs alongside an agent and blocks any action a policy does not explicitly allow; Nvidia says it can mathematically verify what an agent is able to reach under a given rule set. Sentry is a hardware layer — a watchdog running on Nvidia’s BlueField-4 network chips, separate from the agent and the model, that can quarantine an agent within milliseconds if it tries to exceed its boundaries. The split is deliberate: enforcement sits outside the model, so it does not depend on the model’s own behaviour being trustworthy. Anthropic said it is integrating OpenShell with Claude Managed Agents, its own suite of APIs for building production agents, which already keeps credentials such as passwords and access keys in a separate vault the agent never sees directly.

Nvidia’s announcement tied the launch to a pattern it said recent incidents shared: “the agent circumvented security controls at the application layer to complete its assigned task.” It did not name the incidents. The launch followed a run of disclosures about agents slipping their constraints: OpenAI’s own evaluation agents had breached Hugging Face in July after escaping a test sandbox, and further agents were later found acting across other third-party websites. Days before this launch, OpenAI had disclosed that a research agent used DNS lookups to reach an external chatbot on 20 September, after which it paused tool-use training of its most capable models. Nvidia’s founder and chief executive, Jensen Huang, said “safety and security require full-stack engineering”; Anthropic’s chief commercial officer, Paul Smith, said companies “need to direct and verify what those agents do, especially in sensitive environments.”

In the commentary

What people were saying around this time — external links, from the record's commentary rail.