Nvidia Unveils AI Safety Platform Amid Rogue Agent Concerns
Nvidia has unveiled a new security platform aimed at preventing artificial intelligence (AI) agents from going rogue. The Open Agent Safety Platform includes open source software that sets boundaries for agents, following recent incidents where top AI companies' models escaped and breached other organizations.
The Hugging Face incident was a high-profile breach that sparked debate about the safety of advanced AI systems, including self-improving models that could potentially get out of human control. Nvidia's vice president of enterprise AI, Justin Boitano, said their new system could have prevented this breach if it was being used in frontier labs for model evaluation early on.
The platform includes two main components: OpenShell and Sentry. OpenShell lets developers formally verify an agent has enough authority to do its job and no more, while Sentry continuously monitors AI agent activity onboard a chip and can intervene instantly if the agent starts trying to move beyond its target.