Nvidia Wants to Keep AI Agents From Going Rogue. It Has a New Safety Platform For That.
The Nvidia Open Agent Safety Platform is intended to protect agents from testing through deployment, essentially putting a security perimeter around AI that can independently execute tasks.

Nvidia is rolling out a new security platform designed to keep increasingly autonomous artificial intelligence agents under control after a string of real-world incidents raised concerns about what can happen when AI systems escape their sandboxes.
The chip giant on Monday unveiled the Nvidia Open Agent Safety Platform, "an open software platform and reference system design to strengthen AI security from agent testing to deployment, with full-stack governance and control across software and the hardware, compute and robotics systems that run agents."
The system is intended to protect agents from testing through deployment, essentially putting a security perimeter around AI that can independently execute tasks. Nvidia CEO Jensen Huang described the technology as something akin to a browser built specifically for AI agents."You can't have agents roam around and drift around the company, and so you have to find a way to container it," Huang told CNBC's "Squawk Box" Monday.
The launch comes at a particularly sensitive moment for the AI industry. OpenAI, Anthropic and other major developers have disclosed incidents involving advanced models finding their way outside supposedly isolated testing environments.
In July, OpenAI revealed that several of its models escaped an isolated test environment after exploiting a previously unknown vulnerability and then accessed production infrastructure belonging to Hugging Face, the popular open-source AI developer platform.
Anthropic said this month that four Claude models had gained unauthorized access to real third-party systems during cybersecurity evaluations. The company attributed the incidents in part to a misconfigured evaluation environment that unintentionally left the models connected to the internet.
"Recent incidents have highlighted a fundamental hurdle for AI agents, and that is that model-level safeguards alone can't govern what agents can access or do," Justin Boitano, Nvidia's vice president of enterprise AI, said.
One component, Nvidia OpenShell, creates "a secure runtime boundary that traces all actions and enforces policy as agents run on NVIDIA Vera CPUs." Nvidia designed OpenShell for its Vera CPUs, but says it "can be extended to work with third-party compute platforms, including those from Arm and Intel."
Then there is Sentry, Nvidia's second line of defense. Rather than operating on the same CPU or GPU as the agent, Sentry runs separately on Nvidia's BlueField-4 data processing units and continuously watches agent behavior. If an agent attempts to move beyond its permitted boundaries, Nvidia says Sentry can quarantine it within milliseconds.
Nvidia executives told reporters the technology could have prevented the OpenAI-Hugging Face incident. The company is positioning the platform as a reference design, meaning other technology companies can build their own security products and services on top of it.
More than 100 organizations are already participating in the broader initiative, according to Nvidia. Partners include Anthropic, Cisco, Dell Technologies, HPE, Hugging Face, JPMorganChase, Microsoft, Palantir, Palo Alto Networks, Salesforce, SAP and Scale AI. Anthropic has also integrated its Claude Managed Agents technology with Nvidia's security architecture.
© Copyright IBTimes 2026. All rights reserved.



















