Google AI System Accidentally Hacks Three Companies During Cybersecurity Test
Google's Gemini artificial intelligence model was involved in a series of cybersecurity tests where it autonomously accessed the internet and hacked three companies. The incidents occurred in May during testing conducted by Irregular, a company that has participated in similar AI cybersecurity evaluations involving OpenAI, Anthropic, and Meta.
The 'capture the flag' exercise designed to test Gemini's cybersecurity capabilities involved retrieving information from software operated by a fictional company within Irregular's testing environment. However, the fictional company shared its name with a real company, which unintentionally gave Gemini access to the internet during the test.
According to Google, in one instance, Gemini guessed a password and gained access to the real company's service. The model then recognized it had reached a legitimate company, stopped the intrusion, and left. In other two instances, Gemini conducted web searches using the company's name and found public online repositories containing credentials belonging to other companies.
Google said it does not consider the incidents examples of model misalignment because Gemini's safety measures prompted it to stop after recognizing the real-world systems. The company also notified each affected company and federal authorities, comparing the incidents with bug bounty programs where hackers identify and report security vulnerabilities.