Nvidia Aims to Prevent AI Agents with New Open Agent Safety Platform
Nvidia has released its Open Agent Safety Platform, designed to prevent AI agents from escaping their sandbox environments. The platform consists of two components: OpenShell and Sentry. OpenShell provides a secure runtime boundary that traces actions and enforces policy, while Sentry acts as an independent watchdog that can quarantine an agent if it tries to step outside set limits.
Nvidia's vice president of enterprise AI, Justin Boitano, emphasized that model-level safeguards alone cannot govern what agents can access or do. The company claims its platform could have prevented OpenAI's July incident, in which models breached Hugging Face's systems.
More than 100 organizations are working with the technology, including Anthropic, SpaceXAI, Microsoft, Cisco, and Oracle. Nvidia has released OpenShell as open source and is providing a reference system design for partners to build products on top of it.