Two hours inside a production service
The sharpest number in Google DeepMind's 21 July 2026 announcement is a stopwatch reading. "In just 2 hours, the model uncovered remote code execution vulnerabilities in public APIs and found a memory-corruption vulnerability in a sensitive production service," wrote Raluca Ada Popa and Four Flynn, introducing Gemini 3.5 Flash Cyber, a model specialised for security work.
This is not a demonstration against a toy target. Google DeepMind said the model already operates in Google's internal codebases including Chrome, Android, Cloud, Ads, and YouTube, and that on Chrome's production pipeline it showed "significant uplift" over Gemini 3.5 Flash.
Then comes the line that turns a research post into a business story. "3.5 Flash Cyber will be exclusively available to governments and trusted partners via CodeMender soon, expanding over time," the company wrote, describing a "limited-access pilot program." The same day, Google put Gemini 3.6 Flash and Gemini 3.5 Flash-Lite into general availability. Flash Cyber is the only one of the three that you cannot obtain.
The gap is Google's own measurement of what you do not get
Google has published the size of the capability it is holding back. On the V8 JavaScript engine, across fixed invocations, 3.5 Flash Cyber found 55 unique confirmed issues, against 47 for mainline Gemini 3.5 Flash and 36 for Claude Opus 4.6.
Why it matters: that 47 is the practical ceiling for any commercial security product built on a Gemini model today. Google DeepMind said general availability of foundational CodeMender capabilities comes through generally available Gemini models, which is the polite way of saying that the 55 stays inside.
The action for an owner is small and specific. At renewal, ask your scanning and code-review suppliers in writing which Gemini model they call, and keep the answer on file, because a vendor claiming it runs "the latest Gemini" is describing the 47.
The release terms are the product
Access to this model is decided by who you are rather than what you pay. "Governments and trusted partners" is an eligibility test, and Google DeepMind has published no criteria for it, no application route, and no appeal.
For a European business that is a hard limit rather than a budget line. There is no reseller, no framework agreement, and no enterprise tier that converts money into access, which is an unfamiliar shape for buyers used to negotiating scale and support instead of permission.
Google DeepMind was explicit about the reasoning. "Given the dual-use nature of this technology, we have taken an intentional approach to how we deploy 3.5 Flash Cyber," the company wrote, to give "defenders a head start...while mitigating against broader misuse." That logic is sound, and it is also portable: nothing about it stops at security models.
Safety tuning is now subtracting capability from defenders
The quietest detail carries the most weight. Google DeepMind reported that newer competitor model versions refused the Chrome tasks because of safety guardrails, which means that on this evidence the tuning meant to prevent harm is measurably reducing what a defender can do, while the restricted model keeps the capability intact.
The honest caveat: every benchmark here is Google's own, run by Google, with no independent replication, the term "trusted partners" is undefined, and a two-hour result on Google's own services says nothing about your estate. What the numbers do establish is direction, and the direction is that the strongest defensive tooling concentrates inside the small group of organisations that also own the codebases, while everyone else patches on a delay.
Read next: 87 Years Fell to an Answer Anyone Can Check | The Court Case That Prices AI Training



