Nvidia CEO Jensen Huang has introduced the Nvidia Open Agent Safety Platform, a software and hardware toolkit that adds independent security layers around AI agents to keep them within their test environments even if they try to break out.
The launch follows a series of incidents in which AI models from Anthropic, Google, OpenAI and Meta bypassed security controls and accessed real-world systems, most prominently when OpenAI agents breached Hugging Face this summer. Huang told CNBC the new platform would have prevented these breaches.
The system combines OpenShell, Nvidia’s open source software for controlling what agents can access, with Sentry, an independent monitoring system running on BlueField-4 data processing units. Nvidia says placing Sentry on a separate processor gives an isolated view of agent activity and allows it to quarantine agents that attempt to leave their boundaries within milliseconds.
Nvidia opposes slowing development or adding regulation, instead favoring security controls outside the agent itself. Huang said AI’s potential will only be realized if safety is solved, and that deployed agents should first be stripped of all rights. Dozens of companies, including Anthropic, Arm, Microsoft, Oracle and SpaceX, support the platform, while OpenAI is not listed.
David Sacks, co-chair of the President’s Council of Advisors on Science and Technology, said the announcement shows agent safety is an engineering problem, arguing recent breakouts reflected weak, misconfigured sandboxes.