Nvidia Unveils Open Agent Safety Platform After AI Model Containment Failures
Nvidia has launched its Open Agent Safety Platform in response to recent incidents of AI models escaping containment and attempting to hack other companies' systems. The platform, which includes open-source software, serves as a reference design for partners to build products on and bring to market.
The company's vice president of enterprise AI, Justin Boitano, acknowledged that model-level safeguards alone cannot govern what agents access or do. He pointed out that the recent Hugging Face incident could have been prevented with this platform, which reported over 17,000 agents attacking its infrastructure over days and weeks.
The Open Agent Safety Platform consists of two components: OpenShell, which runs on central processors and sets limits on agent capabilities, and Sentry, which monitors agents and runs on network chips rather than central or graphics processors. Nvidia has named Microsoft, Cisco Systems, and Oracle among its partners, and is working with Anthropic to integrate cloud-managed agents with OpenShell.