Two Tiers Instead of One Gate
OpenAI launched Daybreak in May 2026 as a way for vetted security partners to use its most advanced models on real defensive work. On August 10, the company split the program in two. Daybreak Blue opens frontier general-purpose models, including GPT-5.6 Sol, to approved defenders with safeguards adjusted for legitimate security tasks: vulnerability discovery, secure code review, malware analysis, incident response, patch validation. Daybreak Red sits behind a second, tighter vetting layer and gates access to a new model built for a narrower and more dangerous job.
That model is GPT-5.6-Cyber. It is built on GPT-5.6 Sol but trained specifically to reduce refusals on dual-use cybersecurity tasks - the requests a safeguarded model normally declines, like developing a working exploit chain rather than just describing a vulnerability class. OpenAI's own framing, posted the same day, calls this "expanding Daybreak as the cyber defense window narrows" - the company's language for a threat environment it says is moving faster than a single access tier can safely serve.
What 95 Percent Actually Measures
OpenAI grades these models on what it calls the Advanced Cybersecurity Completion Rate - the share of sensitive, dual-use security tasks a model will actually complete rather than refuse. GPT-5.6-Cyber scored 95.0 percent. The prior specialized model, GPT-5.5-Cyber, scored 57.3 percent. The standard, safeguarded GPT-5.6 Sol model - the same one available through Daybreak Blue - scored between 1.5 and 2.0 percent, because its guardrails are built to decline almost all of these requests by default.
Read the three numbers together and the story is not that GPT-5.6-Cyber exists. It is that exploit-relevant task completion jumped from a single-digit floor, to just over half, to near-total in two model generations. A capability that used to require a specialist's sustained attention over days or weeks is now something a single access-gated model call finishes on the first pass in the overwhelming majority of cases.
Not a Benchmark Score - a Patched CVE
The number stopped being abstract when OpenAI used GPT-5.6-Cyber to find two previously unknown vulnerabilities in Chrome's V8 JavaScript engine, chainable to corrupt memory and escape the browser's sandbox. One of them, tracked as CVE-2026-15903, involved a compiler bug where a skipped safety check let an attacker read or write memory inside Chrome's sandbox. OpenAI reported it to Google through coordinated disclosure, and Google has already shipped a fix.
Jared Atkinson, chief technology officer at the security firm SpecterOps, described the practical difference in a comment picked up alongside the announcement: the model "has completed work in under a day that earlier models had not resolved after weeks of" intermittent effort. A working sandbox-escape chain in a browser used by billions of people, found and reported inside a single day, is the concrete case the benchmark score was standing in for.
The Race Is Now Patch Cadence, Not Capability
The headline risk is not that OpenAI built a model that can chain browser vulnerabilities - defenders have always needed that capability, and gating it behind Daybreak Red's vetting is a real control, not a formality. The overlooked risk is what a near-total completion rate on a real-world target means for the timeline everyone else is operating on. If a vetted lab can turn a Chrome zero-day into a working sandbox escape in under a day, the assumption that adversaries need weeks to reach the same result no longer holds as a planning baseline, regardless of whether they have OpenAI's own model or an equivalent capability built elsewhere.
EU organizations already operate under NIS2's incident-reporting and risk-management obligations, and UK organizations follow NCSC guidance on patch timelines; both frameworks were written for a world where discovery-to-exploit took weeks, not a day, and neither currently distinguishes a patch that shipped from a patch that is verified running on every affected endpoint. CVE-2026-15903 got fixed fast because a friendly lab walked it straight to Google - that speed was a courtesy, not a guarantee the next chain gets the same treatment from whoever finds it first. That is the gap for EU and UK security leads to close now: tighten patch-verification SLAs on browser and OS zero-days to match a discovery timeline measured in hours, not the reporting cycle built for one measured in weeks.
Read next: One Open Endpoint Became The Attack's Home Base | Microsoft's Cyber Model Sends the Hard 10% to OpenAI



