Skip to main content
Kaino.dev
Discover
Evals
News
Academics
Insights
Kaino.dev

Discover, evaluate, and compare AI tools, models, and agents.

Explore

  • Discover
  • Evaluations
  • News
  • Academics
  • Insights

Community

  • Twitter
  • YouTube
  • Instagram
Privacy PolicyTerms of Service

© 2026 Kaino.dev. All rights reserved.

Version 1.1.0
Microsoft introduces MAI-Cyber-1-Flash for MDASH security agents · News · Kaino
Microsoft introduces MAI-Cyber-1-Flash for MDASH security agents
Kaino
YesterdayJul 27, 2026, 12:00 AM0 views

Microsoft introduces MAI-Cyber-1-Flash for MDASH security agents

Microsoft says its new MAI-Cyber-1-Flash model, used inside the MDASH agent system, reached 95.95% on the CyberGym benchmark for reproducing software vulnerabilities from code. The company is positioning the model as part of Project Perception, an AI security effort focused on vulnerability management.

flashMicrosoftAI agents

Microsoft introduced MAI-Cyber-1-Flash, a cybersecurity-focused AI model designed to run inside its MDASH agent orchestration system.

A cybersecurity model inside MDASH

In a Microsoft AI post titled “Introducing MAI-Cyber-1-Flash inside MDASH,” the company says the model is part of a broader system for investigating and remediating software vulnerabilities. Microsoft describes MDASH as an agent-and-orchestration environment that can coordinate more than 100 specialized agents, each with distinct roles, tools, prompts and stopping rules.

Microsoft’s Official Blog connects the launch to Project Perception, a security initiative that brings MAI-Cyber-1-Flash into MDASH for software vulnerability management. The company says the combination is intended to help identify, reproduce and address hidden flaws across large codebases.

The important distinction in Microsoft’s description is that MAI-Cyber-1-Flash is the model, while MDASH is the surrounding system. Microsoft says the harness, security context and action space are separated from the model family, which means the underlying model can in principle be swapped or combined with others while keeping the broader workflow intact.

CyberGym result: 95.95%

Microsoft says MDASH using MAI-Cyber-1-Flash plus GPT-5.4 reached 95.95% on CyberGym. The same result is cited by Axios, which reports that Microsoft said the combination reached a 95.95% CyberGym score while competing models scored around 83%.

Microsoft’s own chart lists GPT-5.5 Cyber at 85.6%, with Gemini 3.5, GPT-5.6 Sol and Mythos 5 clustered around 83% to 84%. The Official Microsoft Blog rounds the MDASH result to 96% and says the setup delivered nearly 50% cost savings compared with the current MDASH configuration.

Those figures should be read as benchmark results reported by Microsoft, not as proof that the system will perform at the same level in every production security environment. Cybersecurity work depends heavily on codebase quality, available context, tool access, vulnerability type and deployment constraints.

What CyberGym measures

CyberGym, developed by researchers associated with UC Berkeley, describes its Level 1 benchmark as a test of whether AI agents can reproduce real-world software vulnerabilities. According to CyberGym, agents receive a vulnerability description and an unpatched codebase, then are scored on whether they can produce working proof-of-concept reproductions for target vulnerabilities.

CyberGym says the benchmark includes 1,507 instances from 188 projects. That makes the test more directly related to software security engineering than general coding benchmarks, because it evaluates whether an agent can move from a vulnerability description and source code to a concrete reproduction.

For defenders, that capability matters because reproducing a vulnerability is often a necessary step before confirming severity, prioritizing fixes and validating patches. Microsoft is presenting MAI-Cyber-1-Flash as a way to make that process faster and less expensive inside MDASH.

Why Microsoft is emphasizing AI security

Microsoft’s Official Blog frames Project Perception as part of a broader shift in security operations as attackers and defenders adopt AI. The company says it sees more than 100 trillion security-related observations every day across its products and services, a scale it uses to argue that automated analysis is becoming necessary.

Axios reports that Microsoft is launching new AI agents and a specialized cyber model as part of its effort to counter increasingly AI-enabled threats. The article positions the announcement as both a defensive product move and a response to the changing security landscape.

The near-term question is whether results from CyberGym translate into measurable improvements for enterprise vulnerability management. Microsoft’s benchmark result is unusually high compared with the other systems it lists, but the practical value will depend on how reliably MDASH can operate on messy internal codebases, how it handles false positives, and how safely it can support remediation without introducing new errors.

For now, Microsoft has made a clear technical claim: a specialized cybersecurity model, when placed inside a structured multi-agent system, can outperform more general model configurations on a benchmark designed around real vulnerability reproduction.

Key takeaways
  • 1

    Microsoft introduced MAI Cyber 1 Flash, a cybersecurity focused AI model designed to run inside its MDASH agent orchestration system.

  • 2

    Microsoft describes MDASH as an agent and orchestration environment that can coordinate more than 100 specialized agents, each with distinct roles, tools, prompts and stopping rules.

  • 3

    Microsoft’s Official Blog connects the launch to Project Perception, a security initiative that brings MAI Cyber 1 Flash into MDASH for software vulnerability management.

Continue reading

Latest from Kaino News

Story pulse

Freshness

Yesterday

Views

0

Reading

3 min

Byline

Kainotomic Team

Utilities

Topics

flashMicrosoftAI agents

Sources

Reference material and original reporting used in this story.

Microsoft AI

Published Jul 27, 2026, 12:00 AM

View source