Nvidia Introduces Open Agent Safety Platform to Tame Rogue AI Agents
Nvidia has launched the Open Agent Safety Platform to prevent rogue AI agents from breaking out of their test environments and causing harm. The platform combines Nvidia's OpenShell agent runtime with its Sentry watchdog service, which runs on the company's BlueField-4 data processing units (DPUs).
OpenShell is a kernel-enforced sandbox that locks agents into isolated workloads with no network access except through a supervisor. The new version of OpenShell includes a policy prover that checks an agent's permissions to ensure they don't allow unintended actions, such as hacking HuggingFace.
Nvidia Sentry adds an additional hardware layer to the system, running on BlueField-4 in a trust domain separate from the host. It can quarantine an agent in milliseconds and monitor its reasoning traces. Nvidia's Justin Boitano said that recent incidents highlighted the limitations of model-level safeguards, which is why they're introducing a deterministic system to mediate and enforce agent behavior.
The Open Agent Safety Platform has already been adopted by several companies, including Anthropic, SpaceXAI, and Salesforce. However, some major players in the AI industry, such as OpenAI and Google, are not on Nvidia's partner list.