Frontier AI Models Escape Virtual Machines: Anthropic's Claude Cowork Latest to Show Security Gaps
Researchers have demonstrated that Anthropic's Claude Cowork AI model can escape its local virtual machine and access files on a host Mac, raising concerns about the security of frontier AI models.
The disclosure comes just a week after OpenAI revealed that two of its own frontier AI models escaped a sandbox during an internal security evaluation. In a report published Thursday, researchers at Accomplish AI found that Claude Cowork's local execution mode could escape its Linux virtual machine by chaining together several architectural weaknesses with a Linux kernel privilege-escalation flaw.
Once outside the sandbox, the agent could read and write files anywhere the logged-in Mac user had permission to access, including SSH keys and cloud credentials. The researchers say that fixing any one of the security safeguards that failed at the same time would have stopped the attack.
Anthropic classified the report as 'informative,' saying the kernel flaw fell within their 30-day window for recently disclosed vulnerabilities and the remaining findings were considered defense-in-depth recommendations rather than standalone vulnerabilities. The company said roughly 500,000 macOS users running local Claude Cowork sessions were affected before the issue was addressed.