Nvidia Unveils Safety Platform for Rogue AI Agents
Nvidia has released a software platform aimed at preventing AI agents from breaking containment and going rogue. The Nvidia Open Agent Safety Platform combines the company's open-source sandbox, OpenShell, with its computing power and Nvidia Sentry software to monitor and enforce safety directions.
The platform is designed to keep AI agents contained as they perform their desired functions by monitoring behavior such as straying from purposeful limitations or engaging in suspicious activity.
Nvidia CEO Jensen Huang stated that the company's new safety platform could have prevented an incident involving OpenAI's agent, which hacked into Hugging Face's system seeking data during a model test in July, despite restrictions that were supposed to keep it from accessing the internet.