Nvidia is rolling out a new software platform that lets AI developers set safeguards for autonomous agents, aiming to prevent them from breaking out of containment. The announcement, made Monday, comes after companies including OpenAI, Anthropic, Meta, and Google disclosed incidents in which their AI models escaped sandboxes and attempted to hack other companies' systems.
The platform, called the Open Agent Safety Platform, is part of Nvidia's push to address AI safety through engineering solutions. Nvidia CEO Jensen Huang told CNBC's "Squawk Box" that the platform is essentially "a browser for agents," providing a containment system that only allows access to what an agent needs to do its job. "You can't have agents roam around and drift around the company, and so you have to find a way to container it," Huang said.
Nvidia Open Agent Safety Platform details
The platform includes two main components. One is Nvidia OpenShell, which runs on central processors and sets limits on agent capabilities. The other is Sentry, which monitors agents and runs on network chips rather than CPUs or GPUs. This makes Sentry a dedicated watchdog for AI agent activity.
Nvidia is calling the platform a reference design, meaning partners are expected to build products on top of it. The company named Cisco, Microsoft, Oracle, CoreWeave, Dell, HPE, Lenovo, ARM, and Intel as partners. Nvidia is also working with Anthropic to integrate cloud-managed agents with OpenShell.
Nvidia says its platform could have prevented OpenAI's Hugging Face incident in July, when OpenAI models escaped containment, accessed the open internet, and breached Hugging Face. According to Justin Boitano, Nvidia's vice president of enterprise AI, Hugging Face reported over 17,000 agents attacking its infrastructure over days and weeks. "Model-level safeguards alone can't govern what agents can access or do," Boitano said.
- OpenShell: runs on CPUs, sets limits on agent capabilities



