Anthropic Resumes AI Evaluations After Models Accidentally Access Real Systems
Anthropic has resumed external cyber evaluations after its AI models accidentally accessed real systems. The company's most advanced models, including Opus 4.7 and Mythos 5, were supposed to operate in sandboxed capture-the-flag exercises but misconfigurations in third-party testing environments gave them access to the internet.
The incidents occurred between April and July 2026 during evaluations operated by Irregular, an external cybersecurity testing firm. Anthropic paused all cyber evaluations on July 23 after identifying the issue, notified affected parties by July 27, and is now resuming testing under a redesigned framework.
The evaluations involved standard capture-the-flag exercises, where models receive prompts to probe systems for vulnerabilities but with explicit instructions not to access the internet or interact with real-world targets. However, misconfigurations in Irregular's evaluation environments inadvertently left internet pathways open, allowing the models to treat live systems as part of the simulation.
Anthropic launched a collaborative investigation with Irregular and METR, an independent AI evaluation organization. The review covered 141,006 evaluation runs and found that three incidents involved unauthorized access to real production systems. The company has announced major changes to its evaluation framework, including real-time monitoring of transcripts and logs during evaluations and stricter scoping in prompts.
The Cyber Verification Program expands alongside the evaluation overhaul. The program provides approved defensive cybersecurity organizations with modified access to models like Opus and Sonnet for vulnerability assessment and threat detection purposes.