The Tasalli
Select Language
search
BREAKING NEWS
AI Deep Research · 0 sources Sep 17, 2026 · min read

Microsoft exec called AI scraping ‘the largest theft of labor in human history,’ new unredacted filings reveal

The harshest description of AI data scraping may not have come from a novelist, a photographer or a newspaper editor. According to newly unredacted court filing...

Admin

The Tasalli

Microsoft exec called AI scraping ‘the largest theft of labor in human history,’ new unredacted filings reveal
728 x 90 Header Slot

TL;DR — Quick Summary

**What happened:** Newly unredacted court filings show a Microsoft executive privately described AI scraping as "the largest theft of labor in human history." **Why it matters:** The same record indicates Microsoft and OpenAI built datasets that included paywalled Times content — while warning internally that the practice could gut publishers. **Key takeaway:** The industry's public defence of scraping now sits awkwardly beside its own private words.

Key Facts
**Main Update
** Newly unredacted court filings show a Microsoft executive called AI scraping "the largest theft of labor in human history."
**Impact
** The filings indicate Microsoft and OpenAI scraped paywalled Times content and built datasets from it, with internal warnings it would damage publishers.
**Official Response
** No detailed public statement from either company on these specific filings is captured in the material available at the time of writing.
**Current Status
** The documents are now part of an active legal dispute over how AI models are trained on copyrighted journalism.
**What Next
** Expect the internal language to be argued over in court — and in the wider fight over AI content licensing.

The harshest description of AI data scraping may not have come from a novelist, a photographer or a newspaper editor. According to newly unredacted court filings, it came from inside Microsoft — in the company's own words, about its own practices.

A Microsoft executive privately described AI scraping as "the largest theft of labor in human history," the filings reveal. The same record indicates that both Microsoft and OpenAI scraped paywalled Times content, built datasets from it, and warned internally that doing so would gut publishers.

That gap — between what the industry says in public and writes in private — is now sitting in a court file, where it is far harder to walk back.

What the Unredacted Filings Put on the Record

The documents are notable less for a single sentence than for the combination of them. A senior figure at Microsoft characterises mass scraping as theft of human labour. Training datasets were assembled using paywalled Times journalism. And internal discussion acknowledged the likely commercial damage to the publishers whose work was used.

Read together, the filings suggest the companies understood the consequences of their data practices long before those consequences were publicly litigated.

Why a Private Word Like 'Theft' Carries Unusual Legal Weight

"Theft" is not a neutral word. Used internally by a defendant's own employee, it can be read as evidence of knowledge — the awareness that a practice was contested, not merely debatable.

In litigation, intent and awareness often matter as much as the act itself. The filings do not settle the legal question of whether scraping copyrighted journalism is unlawful. But they may complicate the argument that nobody inside these companies considered the ethics of it.

Paywalled Journalism Was the Raw Material

Paywalls exist to create scarcity. Journalism is expensive to produce, and subscriptions are how newsrooms fund reporting that advertisers no longer pay for.

When that content enters a training dataset, the paywall stops functioning as a boundary. The material still gets read — just not by anyone whose subscription paid for it.

That is the distinction publishers have been trying to make in court: this is not about public web pages, but about work that was deliberately kept behind a paid barrier and used anyway.

Who Actually Feels the Impact

Start with the reporter who spends weeks on an investigation that a chatbot will summarise in a sentence, without a click, a byline or a subscription.

Then consider the smaller newsroom with no legal budget, watching its traffic erode while larger publishers negotiate licensing deals. And the reader, who gets fluent answers built partly from work no one credited or paid for.

The filings describe a systemic concern. The impact, however, lands one newsroom at a time.

How the Dispute Reached This Point

The fight over AI training data has moved in stages: publishers raised objections, licensing talks stalled, lawsuits followed, and both sides entered discovery — the stage where internal documents become evidence.

Discovery is where public messaging and private communication tend to diverge most sharply. Filings are initially sealed or redacted; when they are unsealed or unredacted, language that companies never expected to be read aloud becomes part of the public record.

That is what has happened here. The documents were not written for a judge, a competitor or a reader. They were written for colleagues.

What the Companies Have Said — and What This Record Does Not Show

This report is deliberately narrow. It is based on the filings as described, not on a full reading of a court docket, and it does not attribute any statement to Microsoft or OpenAI beyond what the documents themselves contain.

No detailed public response from either company to these specific filings is captured in the material available. That absence is itself worth noting — not as an admission, but as a gap readers should recognise.

Confirmed Facts — and What Remains Unsettled

On the record: a Microsoft executive described AI scraping as "the largest theft of labor in human history"; both companies scraped paywalled Times content and built datasets from it; internal warnings existed about the damage to publishers.

Not established by these filings: whether any court has ruled on the legality of that scraping, what damages — if any — will be awarded, and how the companies will formally respond to this specific language. Anything beyond the documents is speculation, and should be labelled as such.

The Moat: Why Training Data Is the Real Asset

For readers outside the industry, the business logic is simpler than it sounds. AI models improve with scale — more data, more compute, more users — and each of those reinforces the others.

Microsoft brings cloud infrastructure, enterprise distribution and a deep installed base. OpenAI brings one of the most recognised consumer AI products in the world. Together, they can train larger systems and reach more users than almost any rival.

That is the network effect at work. The harder question the filings raise is what happens when the raw material feeding that moat was never licensed from the people who made it.

The Other Side: Fair Use, Innovation and the Risk of Overreach

AI companies have argued, broadly, that training on publicly available material is transformative use that produces new work rather than substituting for the original. A ruling against them, they warn, could slow research and entrench incumbents who can afford licences.

There is a legitimate case in that. Copyright law has historically allowed limited use of protected work for criticism, teaching and innovation. Overcorrection carries its own costs.

But publishers counter that "publicly available" is not the same as "free to take," and that a search engine indexing a page is not the same as a model absorbing an entire archive. Both arguments will now be tested against internal documents that appear to undercut the neutral framing.

A Wider Pattern: Creators, Coders and Content Farms

Journalism is the visible edge of a much larger dispute. Authors, photographers, illustrators, musicians, translators and software developers have raised parallel objections about their work being absorbed without consent or payment.

The phrase in the filings — theft of labour — frames this as an economic question about who gets paid for creative work, not a technical question about how models learn. That framing may prove more durable than any single ruling.

What Readers, Writers and Investors Should Do Now

Writers and creators: keep records of where your work is published, and read the terms you have already agreed to. Licensing disputes are won on documentation as much as argument.

Publishers: track how your archives are being surfaced and cited, and treat licensing terms as a business decision, not a legal formality.

Investors: litigation and licensing costs are now a structural line item for AI-adjacent businesses, not a one-off risk.

Readers: when a story quotes court filings, look for the filings. Summaries — including this one — compress far more than they reveal.

What Happens Next

The internal language is now part of the case. Expect it to feature in arguments over knowledge, intent and damages, and expect both sides to contest how much weight it deserves.

Beyond the courtroom, three paths are plausible: licensing deals that pay publishers for training data, a ruling that sets a boundary for scraping, or a prolonged stalemate where both litigation and scraping continue. Which one emerges depends on courts, not press releases.

Our Take

The most damaging line in these filings is not the word "theft." It is the implication that the people closest to the decision understood the cost to publishers and proceeded anyway.

That does not make the legal case open and shut. Fair use is a real doctrine, and an over-broad ruling could chill legitimate research. But credibility is a currency, and it is harder to defend a practice as harmless when your own colleagues privately described it as taking someone else's labour.

Whatever the courts decide, the filings have shifted the argument. The question is no longer whether AI companies knew. It is what they decide to do now that everyone can see they did.

Frequently Asked Questions

What do the newly unredacted filings reveal?

They show a Microsoft executive privately describing AI scraping as "the largest theft of labor in human history," alongside indications that Microsoft and OpenAI scraped paywalled Times content, built datasets from it, and warned internally that it would harm publishers.

Has any court ruled that AI scraping is illegal?

Not on the basis of these filings alone. The documents are evidence in an ongoing dispute. No ruling establishing that the scraping in question was unlawful is described in the material available, and any claim otherwise would be speculation.

Why does internal company language matter in a lawsuit?

Because it speaks to awareness. Courts often weigh what a party knew and when. A company's own private characterisation of its conduct is harder to reframe than a public statement drafted by lawyers.

What does this mean for publishers and ordinary creators?

It strengthens the position of anyone whose work was used without a licence, particularly as leverage in negotiations. For smaller creators, the practical value is precedent — evidence that major companies recognised the harm, even if compensation still has to be won.

Are Microsoft and OpenAI the same company?

No. Microsoft is a major investor and cloud partner in OpenAI, and the two have deep commercial ties, but they are separate entities. The filings concern both, which is why the language is being treated as significant.

Written by

Admin