What OpenAI disclosed
OpenAI published a post this week titled "Responding to the next frontier of critical cyber capabilities," describing what its own internal testing found in its next major model, code-named Astra. Axios reported the story on 7 August 2026, the first outside account of the decision. OpenAI said preliminary evaluations showed performance strong enough that the company "cannot rule out" Astra reaching the Critical classification on its Preparedness Framework's cybersecurity track. Astra remains unreleased, and OpenAI said the model played no part in the Hugging Face exploit that circulated in security circles this summer. The company is not announcing a new release date. Instead it is pausing internal work on Astra that does not meet a stricter set of security controls, and slowing the model's overall development timeline until those controls are in place.
OpenAI Confirms the Pause
On 18 August 2026, eleven days after the disclosure above, OpenAI published a follow-up post titled "Pacing model development in an era of cyber-critical capabilities," confirming that the slowdown had become a concrete, dated action. The company said it had paused reinforcement-learning training on its latest deployment-track models for two weeks, and that its largest planned frontier RL run, the training widely understood to sit behind Astra, remains on hold while smaller-scale training and evaluations continue. OpenAI said its expanded monitoring, which checks the model's reasoning at every sampled step and escalates anything concerning to human reviewers within 30 minutes, adds roughly 20 percent to the compute cost of the affected training runs. Chief research officer Jakub Pachocki said it is now necessary to start building tools for coordinating this kind of pacing across AI labs and across countries, and CEO Sam Altman said this was a good time to slow down and that getting safety right matters more than any one company's competitive momentum. OpenAI presented the pause as a decision it made on its own, without waiting for an industry-wide pacing standard to exist, while explicitly asking the rest of the industry to build one.
What a Critical cyber capability actually means
OpenAI's Preparedness Framework, first published in 2023 and updated since, defines the Critical cybersecurity tier as a model that can independently identify and develop functional zero-day exploits, across severity levels, against many hardened real-world systems without a human operator, or that can devise and carry out an entire novel cyberattack against a hardened target from nothing more than a high-level goal. That is a specific, narrow bar, and it sits above the High tier that OpenAI's currently shipping model, GPT-5.6 Sol, was rated at. The distinction matters because a High-tier system still needs a skilled human directing the attack at key steps; a Critical-tier system does not. Astra is the first OpenAI model whose internal testing has come close enough to that line that the company will not say it falls short.
The safeguards OpenAI is now building
OpenAI's response has three parts. Development work on Astra that touches its more advanced capabilities is now confined to isolated, heavily guarded evaluation environments rather than OpenAI's general engineering infrastructure. The company is deploying what it calls universal monitoring across Astra's agentic use: automated systems that watch the model's own reasoning steps in real time and are built to interrupt actions that look misaligned or dangerous before they complete. OpenAI also says it will bring in outside parties, government agencies and independent AI safety institutes, to stress-test Astra's capabilities rather than rely solely on its own evaluations. Each of these is a response to a capability gap, not a policy statement; OpenAI is not claiming Astra is safe, only that it is now testing it as if it might not be.
What this means for EU and UK enterprises
This is the first time a frontier AI lab has publicly slowed a flagship model's release over offensive cyber capability specifically, rather than bias, misuse content or general safety concerns, and that creates a new fact for European and UK risk teams to work with. Under the EU's NIS2 directive and the incoming DORA rules for financial entities, organizations must assess and document the cyber risk their ICT suppliers carry, including AI vendors. Until now, that assessment has had almost nothing concrete to point to: AI labs publish safety commitments, not capability-tier results. OpenAI's admission gives risk and procurement teams an actual artifact, a named threshold with a public result, to put into a vendor questionnaire: has this system been evaluated against a Critical cyber-capability threshold, and what was the finding? UK enterprises reporting to the National Cyber Security Centre and EU entities under NIS2 should expect this kind of disclosure to become a standing line item in AI vendor due diligence, the way SOC 2 reports and ISO 27001 certificates already are. Enterprises with production access to frontier models, particularly anyone running agentic AI against internal networks, should ask their vendor now which capability tier the model they use was rated at, rather than wait for a vendor to disclose it unprompted. OpenAI's own follow-up on 18 August, a confirmed training pause rather than a promise, is itself now the kind of evidence procurement teams should be asking every AI vendor to produce.
Read next: EU Security Teams Have No Say Over OpenAI's Astra | A Model Breach Just Rewrote OpenAI's Safety Rules



