Microsoft MDASH System Outperforms Large Language Models in Cybersecurity Test
Microsoft has developed MDASH, a cybersecurity system that uses over 100 specialized AI agents to detect and fix vulnerabilities in complex codebases. This approach differs from relying on a single large language model, as seen with Anthropic's Claude Mythos Preview and OpenAI's GPT-5.5.
In a recent test, MDASH scored 88.45% on the CyberGym benchmark, surpassing Claude Mythos at 83.1% and GPT-5.5 at 81.8%. By integrating lighter models like MAI-Cyber-1-Flash in July 2026, MDASH's scores still held strong at 96.55%.
The system has already shown its effectiveness during the May 2026 Patch Tuesday, identifying 16 new Windows vulnerabilities, including four critical remote code execution flaws. This technology could potentially be adapted for auditing smart contracts in Solidity, Rust, and Move programming languages.