What the Unsealed Filings Show

A legal brief unsealed on September 17, 2026 in the New York Times' copyright case against Microsoft and OpenAI quotes a Microsoft director describing the companies' AI training practices, in his own words, as theft. Brent Hecht, Microsoft's Director of Applied Science, wrote in a January 2024 internal message that scraping copyrighted content for AI training was an astonishing theft of unprecedented proportions and perhaps the largest theft of labor in human history. The same filing quotes OpenAI's head of ChatGPT, Nick Turley, describing AI products as posing an existential threat to publishers because they are largely substitutive for the original journalism.

Microsoft CEO Satya Nadella's own deposition testimony, also referenced in the filing, states that paywalled content should be licensed by anyone who wants to use it for grounding or training, and that he would have required retraining of models if he had known paywalled material had been scraped without a license.

Why It Matters: This Is the Fair-Use Defense's Weak Point

US fair-use law asks two questions that these filings speak to directly: did the company know its use could cause harm, and does the resulting product substitute for the original in the market. An internal admission from a company's own applied-science director that the practice was theft on an unprecedented scale answers the first question in the plaintiff's favor before a jury ever hears expert testimony. Turley's existential threat language answers the second. Neither statement was written for litigation. Both were written for each other, which is exactly why plaintiffs' lawyers wanted them unsealed.

This is not limited to Microsoft and OpenAI. Every AI company defending a training-data lawsuit on fair-use grounds now has to assume that its own internal Slack messages, emails, and meeting notes carry the same risk, and that its executives' private assessments of their own products can become the plaintiff's best exhibit.

The Scale, in the Companies' Own Numbers

MetricFigure
Drop in click-through rate to nytimes.com after Copilot launchUp to 93 percent
Copies of New York Times, Daily News, and Center for Investigative Reporting works in OpenAI's training dataOver 91,692
Unique news-publisher works in the Project Mango training datasetOver 160,903
Documents from nytimes.com alone in the Common Crawl datasetOver 2,000,000

These are not the plaintiff's estimates. They come from the companies' own training-data logs and traffic analytics, entered into the record because the companies had already counted them internally.

What This Means If You License or Publish Content

If your business earns revenue from content behind a paywall, a subscription gate, or a licensing agreement, this record changes your negotiating position with any AI vendor that has scraped or might scrape your material. You no longer need to prove the vendor knew scraping paywalled content was risky. A company that trains models on your work can be assumed to have internal documentation of that decision, and courts are now willing to unseal it. The practical move is to treat any current or past AI training exposure as a licensing conversation to open now, not a lawsuit to file only after damage is proven.