What shipped, and the benchmark table underneath it

Microsoft announced MAI-Cyber-1-Flash on 27 July, its first artificial-intelligence model built specifically for cybersecurity. The company describes it as a compact, code-heavy model descended from its MAI-Thinking-1 lineage, and it does not run alone. It sits inside MDASH, the harness of more than 100 agents Microsoft uses internally to find and fix software vulnerabilities, which the company says draws on multiple leading models rather than one.

The headline number is 96 percent on CyberGym, a benchmark for finding real vulnerabilities in real code. Microsoft's own published comparison puts Anthropic's Mythos at 83.6 percent, Gemini at 85.6 percent and GPT on its own at 84.4 percent, with a fourth unnamed model at 83.2 percent. The same announcement says the new pairing runs at roughly half the cost of the configuration Microsoft ships today. Satya Nadella framed it as giving customers frontier-grade security at half the cost.

The commercial vehicle is Project Perception, which enters public preview on 3 August inside Microsoft Defender under consumption-based, pay-as-you-go pricing. It coordinates three classes of agent: red-team agents that map attack paths, blue-team agents that investigate and rank risk, and green-team agents that apply fixes. Microsoft has not published rates.

The saving is a routing decision, not a better model

Read the architecture and the 96 percent stops being a capability claim. MAI-Cyber-1-Flash is built to absorb up to 90 percent of the work inside MDASH and to escalate the hardest tenth of tasks to OpenAI's GPT-5.4. The benchmark score is what MDASH achieves running the two together. Microsoft's model is not beating the frontier; it is handling the volume so that the frontier only has to handle the exceptions.

That makes the cost claim more interesting than the accuracy claim. The 50 percent figure is measured against the configuration Microsoft runs today, which the company lists as GPT-5.4, GPT-5.4 mini and GPT-5.3 codex. In other words, the saving was found by displacing OpenAI from the routine nine tenths of the job while keeping OpenAI for the difficult tenth. This is a vendor reducing its own supplier bill and passing part of it on, which is a perfectly good reason to buy, but it is not the same thing as a better detector.

The consequence sits in the exceptions. The hardest tenth of vulnerability work is, by construction, the part a cheap model could not close: the deep logic flaws and the awkward code paths. That tenth now depends on a third-party frontier model remaining available, remaining permitted under your own policies, and remaining priced as it is today. None of those dependencies appear in the product name, and none of them are yours to control.

Consumption pricing changes what a security budget is

The pricing model deserves as much attention as the model. Pay-as-you-go on a system designed to investigate continuously means your security spend stops tracking headcount or seats and starts tracking how much there is to look at. A noisy estate, a large monorepo, a merger that doubles your endpoint count, a bad month of alerts: each of these now has a direct and immediate cost, because the agents respond to signals rather than to a licence count.

For a European finance director this is a forecasting change, not a line-item change. Seat-based security is a fixed cost you approve once a year in euros and forget. Consumption-based agentic security is a variable cost that rises exactly when you are already under pressure, which is the worst possible correlation for a budget. It is also the moment when nobody wants to be the person who throttled the security tool to hit a number.

None of this argues against buying it. It argues for buying it with a cap, an alert and an owner. Ask for the meter definition in writing before the preview opens, agree an internal ceiling that triggers a conversation rather than a cut-off, and make one named person accountable for watching the curve in the first quarter.

The question to settle before 3 August

There is a governance answer that is now out of date in most vendor registers. If you are in scope for NIS2 or for DORA, you are expected to know who processes your security telemetry and your vulnerability findings. A harness of more than 100 agents running multiple leading models means the honest answer involves at least two model vendors, one of which is a competitor of the company selling you the product. Write that down before an auditor asks.

Three practical steps this week. Ask your Microsoft reseller whether Project Perception consumption is billed inside your existing Defender commitment or on a separate meter, and get it in writing before 3 August. Ask which model classes handle escalated findings and whether you can restrict escalation on policy grounds. Then check whether your own internal rules on sending code or vulnerability detail to a third-party model were written on the assumption that your security vendor was a single processor.

The benchmark will be quoted at you for the next year. Treat 96 percent as the property of a pairing, not a product, and price the pairing accordingly.