What happened on launch day

OpenAI planned to publish its GPT-6 Astra announcement at 2pm ET on September 3, 2026. The post went up, was pulled down, reappeared as a broken link when OpenAI's own account tweeted it at 3:32pm, and only became reliably viewable between roughly 4:20 and 5:20pm, after Sam Altman had already reposted the link.

Across that stretch of a few hours, several of the benchmark numbers in the post itself changed, not once but in multiple snapshots, before some of them were reverted to their original values.

The numbers that moved, tracked snapshot by snapshot

Independent monitoring of the page's successive versions, reported by Fortune, recorded the following changes:

MetricEarly published numberLater revisionWhere it stands now
Astra on ARC-AGI-398.6% (embargoed draft)99.99% (live blog)99.99%
Astra hallucination rate4.2% (first snapshot)2% (sixth snapshot)Reverted to 4.2%
GPT-5.6 Sol hallucination rate12.2% (first version)9.4% (later version)Reverted to 12.2%
Claude Fable 5.1, FrontierMath Tier 487.8% (first version)78%, then 83%83%

Two more metrics moved during the same window: an ExploitBench cybersecurity score for GPT-5.6 Sol shifted from 5.5 percent to 11.5 percent before OpenAI said it was investigating a reversion, and a coding benchmark for Astra ticked from 57.7 to 57.9 percent.

OpenAI's explanation, and why it did not fully satisfy researchers

OpenAI told Fortune it cares deeply about getting evaluations right, and that most evaluations carry noise of a few percentage points depending on the exact checkpoint, scaffold, and evaluation run used in reporting. The company described the changes as fixes intended to represent its best estimate of available model performance.

Anka Reuel and Mike Hardy of Stanford's Intelligent Systems Laboratory and Trustworthy AI Lab said the pattern reads as benchmaxxing, tuning toward a good headline number, and pointed to insufficient technical transparency about what exactly changed and why. Vincent Sunn Chen of Snorkel AI was more measured, framing the shifts as normal final launch logistics, but still called for an industry norm of documenting what changes between benchmark revisions and when.

Why the competitor numbers matter more than Astra's own

OpenAI editing its own model's scores upward invites an obvious motive. The more telling detail is that a rival's numbers moved too, in both directions, on the same page, on the same day: Claude Fable 5.1's FrontierMath score dropped nine points then partially recovered, and Claude Fable 5.1 and Opus 5's HealthBench Professional scores were revised upward mid-launch before settling.

A vendor correcting its own claim is routine. A vendor's launch-day page also being the place where a competitor's public score gets revised, more than once, before the industry has finished reading it, is a different kind of event, and it is the one that turns this from a PR footnote into a due-diligence problem.

The decision this changes: what a benchmark card is worth at the moment you read it

Any enterprise buyer who cited Astra's launch-day ARC-AGI-3 or hallucination numbers in a procurement memo this week was, in effect, citing a number that had already changed at least once and might change again. There is no regulatory requirement today that an AI lab version-control its own published evaluation claims the way a financial filing is versioned and dated.

The practical fix is cheap and available now: screenshot or archive.org-snapshot the benchmark page the moment you rely on it, note the exact figure and timestamp in the procurement document, and treat a launch-day comparison table as a claim to be re-verified before a contract is signed, not a fixed input to the decision.