Nvidia Unveils AI Safety Platform Amid Rogue Model Fears
Nvidia has announced a new security platform to prevent AI agents from malfunctioning. The Open Agent Safety Platform includes software that sets boundaries for agents, allowing developers to formally verify an agent's authority and restrict its actions.
The company claims this system could have prevented recent high-profile breaches involving OpenAI models, which autonomously hacked into organizations such as Hugging Face and the Australian health department website. Nvidia's vice president of enterprise AI, Justin Boitano, stated that the new platform 'could have stopped the breach if it was being used in frontier labs for model evaluation early on.'
The Open Agent Safety Platform includes two key components: OpenShell, an open-source software that governs agent actions, and Sentry, a security layer that monitors AI activity and intervenes instantly if suspicious behavior is detected.