Quick Facts
- Microsoft’s MAI-Cyber-1-Flash handles 90% of queries inside its MDASH vulnerability system, routing the hardest 10% to GPT-5.4, cutting costs by about half.
- The combined MDASH system scored 95.95% on the CyberGym benchmark, versus roughly 83% for competing models from Anthropic and OpenAI.
- Project Perception, Microsoft’s agentic defense platform coordinating red, blue, and green team AI agents, enters public preview on August 3.
Microsoft unveiled two new security products on Monday: MAI-Cyber-1-Flash, its first in-house cybersecurity AI model, and Project Perception, an agentic defense platform designed to give enterprises continuous, automated protection. Microsoft CEO Satya Nadella said the system delivers “world-class performance at 50% of the cost of leading models.” Shares rose about 3% on the day.
MAI-Cyber-1-Flash is a compact, code-tuned model derived from Microsoft’s MAI-Thinking-1 line. It was trained on the company’s own exploit and remediation records and reviewed by Microsoft’s AI Red Team, adversarial testing, and an outside assessor before release.
The model sits inside MDASH, Microsoft’s multi-agent vulnerability identification and remediation system. MAI-Cyber-1-Flash carries 90% of all queries — detecting vulnerabilities, patching them, and confirming fixes worked. The remaining 10% of harder tasks route to OpenAI’s GPT-5.4. Microsoft CEO of AI Mustafa Suleyman described GPT-5.4 as “about 10X larger” than the flash model.
That routing split is central to Microsoft’s cost argument. By keeping most work on a smaller, specialized model, the company says it cuts the total cost of running the harness by roughly half compared to relying on frontier models alone.
On the CyberGym benchmark, which covers 1,507 vulnerability reproduction tasks, MDASH running MAI-Cyber-1-Flash scored 95.95%. Microsoft said competing models, including Anthropic’s Mythos and OpenAI’s GPT-5.5-Cyber, scored around 83%. Suleyman said the system also beats Gemini and GPT-5.6 Sol on the same benchmark.
One caveat: the CyberGym results come from Microsoft’s own evaluation. The comparison pits a full tuned agentic system against competitors’ base models, not a controlled model-versus-model test. Suleyman acknowledged the performance comes from the full system. “These are very complicated, long, agentic loops which require storing state, drawing on another database, consulting best practice,” he said.
Enterprise controls inside MDASH include role-based access, tenant isolation, encryption, auditability, and sandboxed execution environments with no internet access.
Project Perception is the second major announcement. The platform coordinates three classes of specialized agents in a continuous loop: red team agents that find vulnerabilities before attackers can, blue team agents that investigate and assess risk, and green team agents that take corrective action and strengthen defenses.
Hayete Gallot, EVP of Microsoft Security, wrote that the platform brings together “signals, context, models and specialized agents into a continuously learning system of defense.” She added that the defining characteristic of next-generation security systems will not be generating more alerts.
The system includes signals and sensors, security context, a coordination harness, and actuators that convert decisions into protection. Microsoft describes it as a continuous learning system that can adapt to changing conditions over time.
David Weston, CVP of Enterprise and OS Security, told Axios that Microsoft still has to “earn the right” to grant its agents more autonomy, signaling a cautious approach to automated decision-making in high-stakes security environments.
Project Perception enters public preview on August 3. MDASH with MAI-Cyber-1-Flash is available now. The announcements were made at an event in San Francisco.
Read more: Microsoft launches AI cybersecurity model, agentic defense platform to cut enterprise security costs
This article was written by an AI agent. Spotted an error? Send a correction and we will fix it.
