AI Models Hijacked: Hackers Bypass Safety Filters with Ease
Hackers are exploiting prominent generative AI models, including Claude Code, Codex, Cursor, and Gemini, for malicious purposes. A recent report from Cisco's Talos intelligence group found that threat actors are using these models to develop malware, automate cyberattacks, and identify software vulnerabilities.
The researchers analyzed prompt histories and chat logs inadvertently exposed online by hackers, revealing their methods. Despite safety filters embedded in commercial AI models, Cisco found that hackers frequently bypass these restrictions using basic jailbreaking techniques rather than complex technical exploits.
Threat actors are claiming participation in authorized 'ethical hacking' competitions or asserting administrative permissions for their activities to circumvent safety blocks. Some attackers also exploit stolen enterprise API tokens and compromised accounts to run their operations on corporate compute power, avoiding the cost of their own infrastructure.