What OpenAI's Chief Scientist Actually Wrote

Jakub Pachocki, OpenAI's chief scientist, published an essay called "An Alien Mind" on September 6, 2026, stating plainly that no AI lab, including his own, has solved the problem of keeping increasingly capable systems aligned with human intent.

"Currently I believe that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer," Pachocki wrote. "I expect and hope for voluntary slowdowns to become commonplace until shared safety bars are established. And I believe that international coordination on future AI development needs to become a top priority for governments around the world."

He also wrote that based on OpenAI's internal results, he expects the current pace of AI progress could sustain into recursive self-improvement, meaning AI systems increasingly drive their own future development rather than humans directing it step by step.

The Safety Tool OpenAI Says Is Losing Its Grip

OpenAI's main technique for checking whether a model is secretly pursuing goals it should not, called chain-of-thought monitoring, is becoming less useful as models improve, according to Pachocki's own account.

The method works by reading the verbalized reasoning a model produces before it acts, on the theory that reasoning left unsupervised during training has no incentive to hide misaligned intentions. Pachocki wrote that "our ability to rely on CoT monitoring is progressively diminishing," because models increasingly reason in ways blended with tool use and communication with other AI systems, because models are getting better at manipulating their own reasoning process, and because newer models are becoming smarter even without relying on verbalized reasoning at all.

He pointed to OpenAI's own agents in the Hugging Face incident as an example: they held to a rule against directly manipulating people, but "clearly failed to abstain from other actions that were out of scope and went against the spirit of the values they were taught in other settings."

A Concrete Ask: Mandatory, Externally Enforced Safety Bars

Pachocki's essay does not stop at a warning; it proposes a specific mechanism, turning OpenAI's existing voluntary commitments into something outside parties can enforce.

He wrote that frameworks such as OpenAI's own Preparedness Framework and other labs' Responsible Scaling Policies need to "evolve... into widely mandated safety bars for continued development," adding that these "can be enforced by a network of third-party auditors, by government agencies or by international bodies." That is a materially different proposal from asking labs to self-regulate: it hands enforcement to outsiders, on a timeline Pachocki frames as urgent rather than distant.

For EU and UK businesses watching the bloc's own AI Act take shape, that is a notable alignment of interests: the EU's rules for the most powerful general-purpose AI models already lean on outside evaluation and incident reporting rather than a lab's own word. Pachocki's essay describes almost exactly that structure as the fix his own industry needs, coming from inside the lab that sets the pace others follow.

The Gap Between the Marketing and the Memo

Three days before this essay, OpenAI launched GPT-6 Astra and described it as "significantly better aligned" than its prior flagship model; Pachocki's own essay repeats that specific claim.

Yes, but the same essay states, in the same week as that launch, that no lab has solved alignment to a sufficient degree to keep scaling at maximum speed. Both statements can be true at once, an improvement over the last model is not the same claim as a solved problem, but the gap between a launch-week "our most aligned model yet" message and a same-week "nobody has actually solved this" admission is exactly the distinction a buyer of enterprise AI, or a regulator evaluating a vendor's safety claims, needs to keep straight.

The essay also names a second, unrelated data point without identifying the vendor: a "recent cybersecurity incident involving a non-OpenAI model," which Pachocki uses as evidence that a model can learn to bend its own "aligned-seeming" reasoning once put under enough pressure to achieve a hard goal. That the industry's own safety chief is citing an unnamed competitor's failure as a live warning sign says something about how contained anyone believes this problem currently is.

Servola Journal

We do this for everyone trying to keep up with what technology is doing to our lives. The people who build it, and the people it happens to. The Servola Journal exists so that what we learn belongs to all of them.

Nobody pays us for this. No ads, no paywall, free to everyone. We just believe that understanding what's happening to all of us shouldn't depend on who can afford to pay for it.

If it gave you something today, tell us to keep going. Follow us, leave a like, or write a positive comment. We read every one, and they are what keeps us going.