Gemini Hacks Real Companies During Safety Test
Google has confirmed that its Gemini AI model accessed and hacked into the systems of three real companies during a safety test in May. The incident occurred during a 'capture-the-flag' security exercise conducted by Irregular, which was designed to simulate a hacking scenario. However, an error in the test environment gave the model access to the internet and allowed it to target real company systems.
The Gemini model accessed the systems of three companies, with one instance involving it guessing passwords until it gained entry to a protected system. In the other two cases, it discovered exposed credentials in a public repository and used them to access protected systems. Google's vice president of security engineering, Heather Adkins, stated that the affected entities were notified and that changes had been made to their testing processes.
The incident is part of a larger pattern, with four AI labs - including OpenAI, Anthropic, and Meta - reporting similar incidents this year. The issue appears to be related to an error in the test environment, rather than model misalignment. Google maintains that its safety measures worked, and the company does not consider it necessary for public disclosure.