OpenAI and Microsoft Knew ChatGPT Was Starting a 'Doom Loop' for the Web
Key takeaways
- Unsealed court documents in the New York Times v OpenAI case show internal warnings about a 'doom loop' where AI scraping degrades the web's content ecosystem
- The doom loop describes AI systems training on web content, then replacing that content in search results, reducing the incentive for original creation
- The documents are part of a major copyright lawsuit alleging OpenAI and Microsoft used NYT articles without permission to train ChatGPT
- Internal documentation using this language significantly weakens the argument that companies were unaware of these risks
Court documents have a way of surfacing things that corporate communications teams would very much prefer stayed internal. The latest batch, unsealed as part of the New York Times' copyright lawsuit against OpenAI and Microsoft, contain a particularly striking detail: the companies' own internal documentation apparently warned that their AI systems were creating a "doom loop" for the web, and they proceeded anyway.
The phrase "doom loop" appears in documents describing a dynamic where AI systems scrape content from the web to train on, then generate outputs that replace the original content in search results and user behaviour, which in turn reduces the incentive for humans to create original content, which then degrades the quality of future training data. It is a feedback loop that, if it runs long enough, hollows out the information ecosystem that the AI systems depend on. The companies knew this. They kept going.
What the Documents Actually Show
The unsealed records relate to The New York Times' lawsuit, which alleges that OpenAI and Microsoft used copyrighted articles without permission to train ChatGPT and related models. The doom loop warning is significant not just as a legal exhibit but as evidence that the potential consequences of large-scale web scraping were understood internally at the time decisions were being made.
The documents also reportedly reference Google, suggesting the companies were aware of the competitive dynamics around web content and AI indexing. The broader picture that emerges is of organisations that identified a structural problem with their approach, documented it, and continued scaling anyway. Whether that constitutes legal liability is a question for the courts. What it represents from a policy perspective is a harder question.
Why the Doom Loop Argument Matters
The doom loop framing is not new among researchers and critics of AI content practices. It has been articulated in various forms since at least 2023, when concerns about AI-generated slop degrading search quality started gaining mainstream attention. What is new is having internal documentation from the companies themselves using the same language.
There is a real tension at the heart of this. Search engines and large language models both depend on a web that is full of high-quality, human-generated content. If AI systems make it less economically viable to produce that content, by substituting AI summaries for original sources, by training on work without compensating creators, and by capturing traffic that used to flow to publishers, then the web they depend on becomes progressively worse. Less diverse, less accurate, less interesting.
This is already visible in some corners of the web. Content farms running AI-generated articles at industrial scale have degraded the signal-to-noise ratio on many search queries. Independent publishers are reporting declining referral traffic as AI-generated summaries intercept the clicks that used to reach them. Whether this constitutes a full loop or just a partial degradation is debated, but the direction of travel is not.
What Happens Next
The New York Times case is one of the most consequential intellectual property disputes in recent tech history. A ruling against OpenAI and Microsoft could reshape how AI companies are allowed to train their models, potentially requiring licensing agreements with publishers at a scale that would significantly change the economics of frontier AI development.
For now, the doom loop documents are most useful as a conversation-starter about accountability. The argument that companies could not have known about these risks becomes much harder to sustain when internal documents use the specific language of the risk. The companies knew. The question the courts will eventually answer is what knowing and proceeding anyway actually means in law.
For anyone who cares about the long-term health of the web as an information resource, these documents are worth paying attention to. Not because the outcome of the lawsuit is certain either way, but because the internal framing reveals something about how these decisions were made. Awareness of a problem is not the same as responsibility for it. But it is not nothing, either.