What OpenAI disclosed

OpenAI published a post this week titled "Responding to the next frontier of critical cyber capabilities," describing what its own internal testing found in its next major model, code-named Astra. Axios reported the story on 7 August 2026, the first outside account of the decision. OpenAI said preliminary evaluations showed performance strong enough that the company "cannot rule out" Astra reaching the Critical classification on its Preparedness Framework's cybersecurity track. Astra remains unreleased, and OpenAI said the model played no part in the Hugging Face exploit that circulated in security circles this summer. The company is not announcing a new release date. Instead it is pausing internal work on Astra that does not meet a stricter set of security controls, and slowing the model's overall development timeline until those controls are in place.

What a Critical cyber capability actually means

OpenAI's Preparedness Framework, first published in 2023 and updated since, defines the Critical cybersecurity tier as a model that can independently identify and develop functional zero-day exploits, across severity levels, against many hardened real-world systems without a human operator, or that can devise and carry out an entire novel cyberattack against a hardened target from nothing more than a high-level goal. That is a specific, narrow bar, and it sits above the High tier that OpenAI's currently shipping model, GPT-5.6 Sol, was rated at. The distinction matters because a High-tier system still needs a skilled human directing the attack at key steps; a Critical-tier system does not. Astra is the first OpenAI model whose internal testing has come close enough to that line that the company will not say it falls short.

The safeguards OpenAI is now building

OpenAI's response has three parts. Development work on Astra that touches its more advanced capabilities is now confined to isolated, heavily guarded evaluation environments rather than OpenAI's general engineering infrastructure. The company is deploying what it calls universal monitoring across Astra's agentic use: automated systems that watch the model's own reasoning steps in real time and are built to interrupt actions that look misaligned or dangerous before they complete. OpenAI also says it will bring in outside parties, government agencies and independent AI safety institutes, to stress-test Astra's capabilities rather than rely solely on its own evaluations. Each of these is a response to a capability gap, not a policy statement; OpenAI is not claiming Astra is safe, only that it is now testing it as if it might not be.

What this means for EU and UK enterprises

This is the first time a frontier AI lab has publicly slowed a flagship model's release over offensive cyber capability specifically, rather than bias, misuse content or general safety concerns, and that creates a new fact for European and UK risk teams to work with. Under the EU's NIS2 directive and the incoming DORA rules for financial entities, organizations must assess and document the cyber risk their ICT suppliers carry, including AI vendors. Until now, that assessment has had almost nothing concrete to point to: AI labs publish safety commitments, not capability-tier results. OpenAI's admission gives risk and procurement teams an actual artifact, a named threshold with a public result, to put into a vendor questionnaire: has this system been evaluated against a Critical cyber-capability threshold, and what was the finding? UK enterprises reporting to the National Cyber Security Centre and EU entities under NIS2 should expect this kind of disclosure to become a standing line item in AI vendor due diligence, the way SOC 2 reports and ISO 27001 certificates already are. Enterprises with production access to frontier models, particularly anyone running agentic AI against internal networks, should ask their vendor now which capability tier the model they use was rated at, rather than wait for a vendor to disclose it unprompted.