Hackers Leverage AI Models to Develop Malware and Automate Cyberattacks
Cisco's Talos intelligence group has made some alarming discoveries regarding the use of advanced generative AI models by hackers. Researchers analyzed exposed chat logs to understand how threat actors bypass safety measures on popular AI tools. The findings reveal that hackers are leveraging top AI models, including Claude Code, Cursor, Gemini, and Codex, to develop malware, automate cyberattacks, and hunt for software vulnerabilities.
The researchers discovered that even after safety filters were built into commercial AI models, hackers rarely needed complex technical tricks to bypass model restrictions. Instead, they relied on simple jailbreaking techniques, such as participating in an authorized 'ethical hacking' competition or asserting administrative permission to perform the work. Some attackers also used stolen enterprise API tokens and compromised accounts to run their operations on corporate compute power.
Nick Biasini, Senior Technical Leader at Cisco Talos, noted that the AI models are in a tough spot because they have to support people who do vulnerability research for a living or red teaming. He expressed hope that there would be more protection from what hackers were asking the models to do, but acknowledged the challenge of restricting bad actors from using AI for their own benefits.