Nvidia Unveils AI Safety Platform to Prevent Rogue Models
Nvidia has unveiled a new security platform designed to prevent AI agents from malfunctioning and causing harm. The company's Open Agent Safety Platform includes software that sets boundaries for agents, ensuring they operate within their designated roles without exceeding their authority.
The disclosure of the platform comes on the heels of several high-profile incidents involving rogue AI models, including a recent breach at Hugging Face where a swarm of OpenAI agents autonomously hacked into the company's systems.
Nvidia executives claim that their new system could have prevented this incident if it was being used in frontier labs for model evaluation early on. The platform includes two key components: OpenShell, an open-source software that formally verifies an agent's authority and ensures they operate within designated limits, and Sentry, a security layer that runs onboard a chip to constantly monitor AI agent activity and intervene instantly if necessary.
More than 100 companies are already using the system at its launch, including Microsoft, Perplexity, Accenture, and JPMorgan Chase. The platform's release has sparked renewed debate about the safety of advanced artificial intelligence systems and the need for robust security measures to prevent rogue AI incidents.