A Number, Not Just Another Warning

On September 8, Anthropic's own Alignment Science Lead put a specific number on the risk his employer builds toward every day. Evan Hubinger, replying to a departing colleague, wrote that he personally sees a greater than 10 percent chance that AI causes human extinction within the next decade.

The colleague was Jacob Coxon, a 27-year-old researcher who spent three years on pretraining, first at OpenAI on its GPT-4o model, then at Anthropic. Coxon resigned late Tuesday and posted that both labs are "racing straight to self-improving superintelligence and gambling with our lives," adding that the people building the technology "earnestly believe it could kill us all by the end of the decade." Hubinger called him correct and said the risk he has in mind comes specifically from recursive self-improvement, not from the commercial models Anthropic sells today.

Why an Insider's Number Reads Differently

Public estimates of catastrophic AI risk are not new; outside researchers and campaigners have published survey numbers for years. What changed on September 8 is who said it: the person Anthropic itself put in charge of alignment science, on the record, with a figure attached rather than a mood.

Anthropic had not issued a company statement addressing the number by the time this was reported, and Hubinger was explicit that it reflects his personal view, not an official company position. That distinction will not survive contact with a procurement spreadsheet. A named risk owner's number, however personal, is the kind of evidence a compliance team can actually write down.

The Compliance Angle the Coverage Missed

Article 55 of the EU AI Act has required providers of general-purpose AI models with systemic risk to evaluate and document those risks since August 2, 2025, and the AI Office gained the formal power to assess compliance on August 2, 2026, a little over a month before Hubinger's post.

The obligations are specific: standardised model evaluation including adversarial testing, assessment and mitigation of systemic risk across the EU market, reporting of serious incidents to the AI Office, and adequate cybersecurity for the model itself. Providers whose training compute crosses the Act's systemic-risk threshold, the bracket frontier labs like Anthropic sit in, are exactly who this article was written for. A stated risk figure from the person who leads that provider's own alignment work is precisely the kind of foreseeable-risk detail a regulator or an auditor will now go looking for.

What Changes for the Team Buying the Tool

Nothing changes in the product today. No model was withdrawn, no feature was disabled, and Hubinger's own clarification tied his estimate to a future scenario, not to the assistant currently answering your team's questions.

What should change is the paperwork. Vendor risk registers that currently say something like "AI safety concerns, monitored" can now cite a sourced, dated, on-record figure instead. And any team relying on a single frontier lab for a core workflow should treat that dependency as a named concentration risk, not a footnote, given that the number came from inside the vendor rather than from a critic outside it.