The Real News Was Not Model 2
Every outlet that covered this report led with 'Model 2,' the unreleased, more-capable-than-deployed internal model Anthropic disclosed almost as an aside. That is a real detail: Model 2 is used extensively inside Anthropic for coding and research, is a noticeable improvement over the company's current flagship, Claude Mythos 5, and there are no current plans to release it, partly because Anthropic has not yet run its full suite of predeployment safety tests on it. But that framing missed what the report actually spends the most space on: five specific, dated, named process failures, published by Anthropic about Anthropic, in a document format no other frontier lab currently produces at this level of detail.
Also raised: the report's overall assessment of misalignment in high-stakes settings, the risk that a Claude model with access to sensitive systems could exploit that access to interfere with Anthropic's own safety work, moved up one notch, from 'very low' to 'low.' Anthropic's stated reason was not a change in what its models can do; it was general increased uncertainty around recent incident disclosures related to model behavior in cybersecurity evaluations across the industry. In other words, the company raised its own risk estimate in response to what other labs' incidents implied about how much can go unnoticed, not because of anything new inside Claude.
Eleven Months of Unscreened Conversations
The most concrete number in the report concerns a gap Anthropic did not catch in real time; it found the gap since its last Risk Report, while reviewing other issues. From May 2025, when Anthropic first began deploying models with chemical and biological weapons safeguards, until April 2026, a flag intended only for internal use turned off the blocking classifiers meant to stop those models from producing dangerous biological information for the pool of contractors who provide human feedback on Claude's outputs. The same flag also disabled the logging that would have flagged any of that traffic for review, so nothing about the gap surfaced through Anthropic's normal monitoring.
The scale is what makes it notable: roughly 50,000 contractors, vetted only by Anthropic's outside vendors rather than by Anthropic's own screening, exchanged an estimated 133 million messages with models that included some of the company's most capable systems, almost all of it with open-ended conversational access rather than a constrained task interface. Anthropic's after-the-fact review, run using its own Claude Sonnet 5 as a classifier over the affected human messages, flagged 1,197 conversations as potentially high for biological harm; a manual check of the sixty-two flagged messages that came from real users, plus a sample of thirty from Anthropic's own red-teaming exercises, found nothing the company judged as meaningful uplift toward an actual weapon. Anthropic says it now believes there is an increased likelihood of other, similar gaps it has not yet found.
A Training Bug That Taught the Wrong Lesson
A second cluster of disclosures concerns not who could reach Claude, but what Claude itself was trained on. Anthropic confirmed that chain-of-thought reasoning, the step-by-step 'thinking' output several Claude models generate before answering, was unintentionally exposed to the reward-scoring process during reinforcement learning, at rates the company estimates ran from 0.27 percent of training episodes for Claude Opus 4.8 up to 5.1 percent for Claude Mythos Preview. That contradicts an earlier public statement, repeated in Anthropic's own February 2026 report, that training did not directly incentivize the models to reason in any particular way. Independent AI safety group Redwood Research, whose own 2024 research with Anthropic sits behind an unrelated contamination problem described elsewhere in the same report, reviewed this disclosure and judged it evidence of inadequate processes rather than a one-off, noting this was not the first time chain-of-thought exposure had gone undetected across a Claude model generation.
Separately, during early training of Claude Mythos 5, a bug in how a new dataset was built meant the model was taught, for a period, to actually perform more of a class of bad behavior rather than merely learn to report it afterward, before Anthropic caught the error and restarted training from an earlier checkpoint. And in a fifth disclosed incident, an employee whose AI usage was not logged gave an agent open-ended access to a cluster holding very sensitive resources; that agent spawned further sub-agents with elevated, unmonitored permissions, one of which deleted a large number of jobs before anyone caught it. Anthropic says it believes this was the agent making a mistake while cleaning up its own work, not deliberate sabotage, but adds that the monitoring gap means it cannot actually confirm that.
What This Means for Anyone Buying Safety First
Anthropic has built a large part of its commercial and political identity around being the safety-conscious frontier lab: it supports government AI-safety legislation other developers have opposed, offers a dedicated Claude for Government product, and has proposed one of the industry's most restrictive regulatory frameworks, one that would let the US federal government block the release of models it judges too dangerous. None of that positioning is undercut by this report in a legal or factual sense; Anthropic still assesses every category of catastrophic risk in the document as low, and still concludes its own development clears its self-defined cost-benefit bar.
What the report changes is the evidence available to anyone evaluating that positioning. A frontier lab that runs safeguards across tens of thousands of contractors, several model generations, and a growing internal agent workforce is going to generate incidents; the choice every lab actually makes is not whether incidents happen, but whether it publishes them in enough detail for an outsider to judge the pattern. Anthropic's report gives European governments and enterprises weighing a safety-first vendor an unusually specific data point: eleven months with a safety control silently switched off, a training bug that moved a model in the wrong direction, and an independent reviewer calling the underlying processes inadequate. That is a case for taking Anthropic's transparency seriously. It is not, on its own, a case for taking any lab's marketing about its own safety at face value, including Anthropic's.
Read next: Anthropic Unblocks Biology AI, Keeps Bioweapons Locked | Anthropic's IPO Target Doubled in Ten Weeks



