Microsoft Exec Called AI Scraping the Largest Theft of Labour in Human History
Key takeaways
- Unredacted court filings reveal a Microsoft executive called AI training data scraping the largest theft of labour in human history
- Microsoft is one of OpenAI's largest investors and has integrated generative AI across its product suite
- Major AI companies including OpenAI and Google have begun striking licensing deals with publishers, but individual creators have no equivalent recourse
- EU regulators have been more aggressive than US counterparts in examining AI training data practices under the AI Act
A new batch of unredacted court filings has surfaced a quote that is difficult to look away from. A Microsoft executive, in internal communications now part of ongoing litigation, described AI training data scraping as the largest theft of labour in human history. The quote is striking not because it is an unfamiliar argument, writers, artists, and coders have been making versions of it for years, but because of who is saying it, and in what context.
Microsoft is one of the largest investors in OpenAI, has integrated generative AI across its entire product suite, and is itself a beneficiary of models trained on scraped content. For a senior executive inside that organisation to use the phrase largest theft of labour in human history in internal communications suggests a level of internal friction around this question that the company's public statements have never fully revealed.
The Case Being Fought
The filings are part of broader ongoing litigation around AI training data, the details of which span multiple cases in US courts. The core dispute, repeated across dozens of lawsuits, is whether scraping publicly available text, code, and images from the internet to train AI models constitutes copyright infringement, unfair use, or something that existing law simply did not anticipate and has not yet caught up with.
Courts have been moving slowly and inconsistently. Some early rulings have favoured AI companies, finding that training on publicly accessible data falls within fair use principles. Others have been more sympathetic to plaintiffs, particularly in cases involving more direct copying or where the model's outputs are closely derivative of specific works. The landscape remains genuinely unsettled, and these filings add a new layer of complexity: internal acknowledgement from within the AI industry that the ethical case against scraping has real merit.
The phrase theft of labour is also specifically interesting as a framing. Copyright law protects the expression of ideas, not the labour that went into producing them. If Microsoft's executive is framing this as a labour issue rather than a copyright issue, that maps onto a different and in some ways harder legal and moral problem. You can argue about whether a particular piece of text is copyrighted. It is harder to argue that the person who wrote it should not receive any recognition or compensation when their work forms part of the training corpus for a commercial product generating billions of dollars in revenue.
What This Means for the Industry
The disclosure matters because it complicates the public positioning of major AI companies. The standard industry line has been some variation of: training on publicly available data is legal, widely practiced, and consistent with how humans learn. That argument has genuine defenders, and it is not without merit. But when senior people inside those companies are privately describing the practice in dramatically harsher terms, the gap between public and private positions becomes harder to ignore.
For creators, the practical question is what changes as a result. Several AI companies including OpenAI and Google have struck licensing deals with major publishers, news organisations, and content platforms, paying for access to training data rather than scraping it. But these deals typically cover established, well-resourced organisations. The millions of individual writers, photographers, illustrators, and developers whose work may have contributed to training datasets have no such recourse at the moment.
Regulators in the European Union have been more aggressive than their US counterparts in examining training data practices under the AI Act and existing copyright frameworks. The UK government has been consulting on whether to allow text and data mining for AI without a rights-holder opt-in, a proposal that drew fierce opposition from the creative industries.
The unredacted filing does not change any of that overnight. But it is the kind of document that tends to resurface in future legal proceedings, regulatory debates, and public conversations. When you want to argue that an industry knew its practices were ethically problematic and proceeded anyway, a senior executive calling it the largest theft of labour in human history is a fairly useful exhibit.