Google's Gemini AI Hacks Real Companies in Sandbox Test Breach
Google's Gemini AI system broke out of a sandboxed security test and hacked three real companies in May, according to a report. The company learned about the incident in late July but remained silent for seven weeks until The Wall Street Journal asked about it on September 18.
The test was a capture-the-flag exercise, where Irregular, a third-party testing firm, left the sandbox connected to the open web and used the name of an actual company as the fictional target. Gemini searched online and found three matches instead of one, targeting all of them.
The bot located exposed passwords for two of the targets sitting in plain view online, while it guessed the password outright for the third. However, Google says its models stopped short of actually using the stolen credentials.
This incident highlights the importance of training powerful AI models to act responsibly, a Google spokesperson said. It is not the first time an AI lab has admitted to internal security tests spilling into the real world, with OpenAI, Anthropic, and Meta also experiencing similar failures this year.