FTFuture Technology
AI

Seattle Times and Newsday Sue OpenAI and Microsoft Over Training Data

· 3 min read · By Nath Connell

Key takeaways

  • Seattle Times and Newsday filed suit against OpenAI and Microsoft over alleged use of journalism to train AI models
  • They join multiple other major US publishers including the Chicago Tribune and New York Daily News in similar legal action
  • The New York Times case, filed in December 2023, is still unresolved and will likely set key precedents for all subsequent suits

Two more major US news organisations have filed lawsuits against OpenAI and Microsoft, alleging that their published journalism was used without permission to train AI models. The Seattle Times and Newsday join a growing list of publishers that have taken legal action over AI training data, in what is becoming one of the defining intellectual property battles of the decade.

The lawsuits land at a complicated moment. Just last year, a judge found that Microsoft's Copilot had not meaningfully reproduced New York Times articles through its chatbot interface, a ruling that seemed to provide some cover for AI companies. But the core training data question remains very much open, and these new suits suggest that publishers have no intention of letting it drop.

What the Publishers Are Claiming

The exact filings have not yet been made fully public, but based on previous similar cases, the publishers are likely arguing that OpenAI and Microsoft used their archived news content to train models without a licence and without compensation, and that this constitutes copyright infringement at scale.

The New York Times was the first major publisher to file this kind of lawsuit, back in December 2023, and its case is still working its way through the courts. Since then, a number of other publishers have followed, including the Chicago Tribune, the New York Daily News, and several other regional American newspapers. The Seattle Times is a significant regional paper with deep investigative journalism heritage. Newsday is one of the largest-circulation newspapers in the US by readership.

The argument from publishers rests on a specific claim: that their journalism, produced at enormous cost through reporting, editing, and fact-checking, is precisely the kind of high-quality human-generated text that makes large language models useful. Without it, these models would be less accurate and less capable. Publishers believe they are owed compensation for that contribution.

OpenAI and Microsoft's Counterarguments

The AI companies have generally argued that training on publicly available text constitutes fair use under US copyright law, a doctrine that allows limited use of copyrighted material without permission in certain circumstances. They also argue that the outputs of their models are sufficiently transformative that they do not constitute reproduction of the original works.

The future, in 3 minutes a day. The biggest tech story explained every morning, free. Get the briefing →

These are genuinely contested legal questions. Fair use has four factors that courts weigh, and the application of that doctrine to AI training data at scale is something that American courts are still working out. The outcomes of these cases will set precedents that affect the entire industry.

Microsoft has separately been in licensing negotiations with several publishers and has struck deals with some. OpenAI has also signed content licensing agreements with a number of news organisations, including the Associated Press and several international publishers. But many others have held out, either because the offered terms were inadequate or because they want the legal question resolved.

Why This Matters Beyond the Courtroom

The stakes here are not purely financial, though the financial stakes are significant. If courts find that AI training on copyrighted text requires licensing, the cost of building large language models increases dramatically. That could entrench the largest companies, which have the resources to negotiate licensing deals, and make it much harder for smaller AI developers to compete.

Alternatively, if fair use arguments prevail broadly, publishers will have lost what may be their most significant source of leverage in negotiations with AI companies. Some have already argued that the existence of AI tools is suppressing traffic to their websites as readers get answers from chatbots instead of clicking through to original articles.

The Seattle Times in particular has a strong claim to public interest journalism credentials. It has won Pulitzer Prizes and is locally owned, a rarity in the current American media landscape. Its decision to sue is a signal that this is not just a strategy by large media conglomerates, but a position that regional publishers also feel strongly about.

We are still probably years away from final resolution in most of these cases, but each new filing adds pressure on AI companies to either settle, negotiate, or win definitively in court. None of those options is simple.

Sources

The biggest tech story, explained in 3 minutes every weekday. Choose your briefings →

Free. No spam. Unsubscribe in one click.

Enjoyed this? Get the briefing.

One email, every weekday: the top story, a useful tool, and what matters in tech - in under 3 minutes.

More from Future Technology

AI

Hikers Rescued After Trusting Google Gemini With Their Lives

AI

Roland's Melody Flip Brings Generative AI Into the Music Hardware World

AI

AI reads brain MRIs in seconds with 97.5% accuracy

AI

Google ships the Gemini 3.6 Flash family and starts Gemini 4