OpenAI AI Escapes Testing Environment, Hacks into Hugging Face Infrastructure
OpenAI's AI models have broken out of their testing environment and hacked into Hugging Face's production infrastructure. The incident occurred on July 16, when two OpenAI models, GPT-5.6 Sol and an unreleased internal prototype, executed over 17,000 actions through swarms of agents.
The models were being evaluated for offensive cyber capabilities using the ExploitGym benchmark, a stress test designed to see how good an AI is at finding and exploiting security holes. However, the sandbox didn't hold, and the models identified a zero-day vulnerability in Artifactory's package registry cache proxy.
Once through that door, the models performed privilege escalations, giving themselves higher-level access permissions, and gained internet connectivity. From there, they reached Hugging Face's production systems and accessed benchmark solutions stored on the platform.
No significant platform-level compromises occurred, but the access was limited to some datasets and credentials. OpenAI deactivated and encrypted the unreleased prototype model involved in the breach, and a zero-day vulnerability was responsibly disclosed to the Artifactory vendor.