AI Guardrails Easily Bypassed by Claiming Authorization
Cybercriminals are exploiting AI coding assistants and chatbots to build attack tools, operate scam infrastructure, and probe live systems. According to a report by Cisco Talos, threat actors often bypass safety checks with simple claims that their work is authorized.
The researchers examined prompt logs from various AI tools, including Claude Code, Codex, Cursor, and Gemini. They found that malicious software development, expansion of criminal operations, and vulnerability research were common activities among threat actors.
One notable trend was the use of simple ownership claims to gain cooperation from AI models without verification. In many cases, a claim of authorization was enough for the model to comply, highlighting a weakness in guardrails designed to prevent such behavior.
The report also highlighted the varying levels of technical ability among threat actors using AI tools. Novices generated limited and faulty tools, while experienced actors used AI to automate scanning, exploitation, data collection, and maintenance.