A Fourth Masthead Joins the Same Case
The Seattle Times and Newsday filed suit against OpenAI and Microsoft on September 4, 2026, in the US District Court for the Southern District of New York. Both papers allege the companies scraped their websites, paywalled articles included, and folded that reporting into the datasets that train and run ChatGPT, Microsoft Copilot, and Bing's AI features. They are not asking only for damages. Their complaint asks the court to order destruction of any AI model or training dataset that incorporates their work, the same remedy the New York Times sought when it opened this line of litigation in December 2023.
The case joins a docket that already runs to dozens of publishers, authors, and image libraries suing the same handful of AI companies over the same underlying claim: that training on copyrighted material without a license is infringement, not fair use. The Seattle Times and Newsday are not filing something new so much as adding weight to a case that has been building for nearly three years and is now consolidated as NYT v. Microsoft et al.
The Number Microsoft Put on the Record
In discovery for that consolidated case, Microsoft handed the news plaintiffs' own expert roughly 8.2 million Copilot conversation logs. The company says those logs were not a random sample of Copilot usage; they were filtered specifically because they contained keywords tied to the plaintiff newspapers' websites, meaning they were the conversations most likely to contain a match. Even inside that stacked sample, the expert found only 59,545 conversations, under one percent of the total, that reproduced 16 or more consecutive words from a plaintiff's article. A parallel check across the same logs found 24 matching responses for book content, a rate Microsoft's filing puts at 0.00029 percent.
Microsoft's argument follows directly from the numbers: if a tool built specifically to surface likely infringement still finds it in under one case in a hundred, reproduction is not the systemic behavior the plaintiffs describe, and courts weighing fair use should treat these as rare, not representative.
Why the Same Number Does Not Settle the Case
The filing is a genuine data point in a fight usually fought with rhetoric, and that is worth taking seriously on its own. But a pre-filtered sample built to find matches sets a floor on how often reproduction happens, not a ceiling; a broader, unfiltered sample of all Copilot traffic could show the true rate is lower still, or the filter could have missed paraphrased reproduction that never met the 16-word threshold Microsoft's expert used. The newspapers are not only alleging verbatim copying. They also argue that Copilot and ChatGPT can closely paraphrase their reporting and answer reader questions in ways that remove any reason to visit the original article or pay for a subscription, a harm that a word-match count does not measure at all.
The practical fight is not really about whether reproduction is common. It is about whether a small, provable rate of verbatim copying is enough to justify a remedy as absolute as ordering a company to destroy a trained model, when isolating one plaintiff's text out of a model with hundreds of billions of parameters is not straightforward even if a court grants the order.
The Part That Matters Outside the US Courtroom
European readers should not look to Brussels for the answer to the reproduction question this case is fighting over. The EU AI Act's Article 53 transparency duty, already in force for general-purpose AI providers, requires a summary of the categories of content used in training. It does not require, and will not produce, a reproduction-rate audit like the one Microsoft was forced to run for a US court. That means the practical benchmark for how much verbatim copying a licensing deal or an indemnity clause needs to cover is being set case by case in US litigation, not by EU regulation.
Any European company buying Copilot, ChatGPT Enterprise, or a similarly trained tool for a business that produces its own copyrighted content, journalism, technical documentation, or code, should expect vendor contracts to start citing numbers like Microsoft's 59,545-out-of-8.2-million figure directly, because it is now the closest thing the industry has to an evidenced answer instead of a claimed one.
Servola Journal
We do this for everyone trying to keep up with what technology is doing to our lives. The people who build it, and the people it happens to. The Servola Journal exists so that what we learn belongs to all of them.
Nobody pays us for this. No ads, no paywall, free to everyone. We just believe that understanding what's happening to all of us shouldn't depend on who can afford to pay for it.
If it gave you something today, tell us to keep going. Follow us, leave a like, or write a positive comment. We read every one, and they are what keeps us going.
Read next: OpenAI Wants Its Own Incident Clock. Brussels Started One First. | A Security Badge Told The Pentagon Nothing About What Grok Would Generate



