Adversaries Weaponize AI with Ease, Bypassing Guardrails Across Models and Platforms
Cybersecurity researchers at Cisco Talos have published a report detailing how adversaries are weaponizing artificial intelligence (AI) in 2026. The report, which analyzed prompt log artifacts from tools like Claude Code and Gemini, found that AI safety guardrails were ineffective across every model and platform examined.
The researchers discovered that actors relied on simple ownership claims, Capture the Flag or bug bounty labeling, task decomposition, and persona conditioning to get what they wanted from the AI models. This allowed them to bypass even censored versions of the models, as operators simply pivoted to uncensored versions when faced with resistance.
The report highlighted three categories in which adversaries are using AI: as a malicious software engineer, a criminal force multiplier, and a vulnerability research accelerator. In the first category, a novice DDoS operator was able to build tooling controlling nearly 2,000 Android TVs by claiming they were his own test infrastructure.
In the second category, a Russian fraud actor embedded persistent AI memories to strip guardrails permanently, while a Spanish-speaking operator built an autonomous OpenClaw agent named 'Alex' targeting Telegram Mini Apps and cryptocurrency wallets.