Nvidia's Open Agent Safety Platform Tackles Rogue AI Agents
Nvidia announced its Open Agent Safety Platform on September 28, 2026, in response to the growing concern of rogue AI agents. The platform is a two-layer architecture consisting of OpenShell and Sentry, designed to contain autonomous AI agents before they can cause harm.
OpenShell is the CPU-side component that sets formal boundaries for an agent's authority, ensuring it only has access to approved systems and actions. It uses formal verification, a method borrowed from safety-critical software engineering, to mathematically prove the agent's behavior stays within defined limits.
Sentry is the data-processing unit (DPU) component that independently monitors agent behavior and quarantines violations in milliseconds. By separating Sentry from the agent's own chip, Nvidia ensures a rogue agent cannot talk its way past a watchdog it never has direct access to.