Policy

Microsoft Exec Called AI Scraping 'Largest Theft of Labor' Ever

Unsealed court documents reveal internal warnings that training AI on news content would devastate publishers and create a self-destructive 'doom loop.'

Omega Editorial· September 17, 2026· 4 min read

Internal warnings contradicted public stance

Newly unsealed court documents in the copyright lawsuit led by The New York Times have exposed internal communications showing Microsoft and OpenAI executives privately acknowledged that scraping news content for AI training constituted massive-scale theft—even as both companies publicly defended the practice as fair use.

Microsoft Director of Applied Science Brent Hecht described the practice in internal documents as "an astonishing theft of unprecedented proportions" and potentially the "largest theft of labor in human history," according to the unsealed motion for summary judgment filed by news organizations. Hecht also wrote that widespread news scraping made "a complete mockery of the idea of 'fair use,'" directly contradicting the legal defense both companies have mounted.

At OpenAI, ChatGPT head Nick Turley warned internally that publishers would face an "existential threat" from commercial products trained on their content that could substitute for the original news sources. The details were first reported by Ars Technica.

The 'doom loop' prediction

Internal Microsoft documents described a "doom loop" scenario where AI products would undermine the very content suppliers they depend on. One document stated: "It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its 'content supply chain.'"

The prediction appears accurate. Microsoft's own data showed click-through rate drops of 83–93 percent for some news plaintiffs and 51–94 percent for others. OpenAI's Turley acknowledged that chatbots are "largely substitutive, period" and predicted they "will get more and more substitutive as they get better."

An OpenAI software engineer noted in internal messages that "no matter how prominently we show the links, users won't click." Turley agreed there is "no good reason to click" when chatbots provide the information directly.

Why it matters

The unsealed documents undercut the fair use defense that has been central to AI companies' legal strategy. If courts determine that chatbots substitute for news sites in their own markets while reproducing content verbatim, it could establish that licensing agreements are legally required—fundamentally changing the economics of AI development. The case also highlights a structural problem: while the AI industry collectively benefits from sustaining news production, individual companies have incentives to free-ride on content, creating what news plaintiffs call a "prisoners' dilemma" that only legal intervention can resolve.

Evidence of verbatim reproduction

News organizations documented multiple methods by which chatbots reproduced their articles. Beyond early techniques of repeatedly asking "what's the next line," they found that requesting summaries, bullet points, or asking chatbots to "rate the bias" of articles would generate extensive verbatim excerpts. Chatbots also reproduced content when users asked them to select articles from a site's homepage.

Microsoft CEO Satya Nadella testified under oath that AI companies shouldn't violate news sites' terms of use by circumventing paywalls. However, internal OpenAI messages showed that when a staffer informed President Greg Brockman that "a hack" was found for OpenAI crawlers to bypass the New York Times paywall, Brockman replied, "Ah, nice."

Companies distance themselves from internal warnings

A Microsoft spokesperson told Ars Technica that Hecht's documents "reflect one employee's individual perspective, are not a legal analysis, and do not represent the company's views." The spokesperson defended Microsoft's AI products as transformative fair use that don't substitute for news sites.

Steven Lieberman, counsel for the New York Daily News and seven sister papers, disagreed. "The evidence revealed here for the first time shows that OpenAI and Microsoft knew that what they were doing was wrong," he told Ars Technica. "Throughout this case Defendants insisted that these documents be treated as confidential so that the public could not see them."

The unsealed motion also alleged that Microsoft violated industry norms by selling a dataset purchased for Bing as training data for OpenAI without consulting news organizations, and that OpenAI obtained a New York Times dataset containing 1.8 million articles from a third party despite knowing it was bound to an agreement prohibiting commercial use.

The details were first reported by Ars Technica.

#copyright#ai training data#microsoft#openai#news publishers#fair use

This is an original analysis by the Omega editorial team. Source reporting: AI Watch.

Want systems like this working for your business?

Book a Call

More in Policy

Policy· 3 min read

Microsoft Executive Called AI Training 'Largest Theft of Labor'

Internal comments surface in unsealed court documents from the New York Times copyright lawsuit against Microsoft and OpenAI.

Via AI Watch · Sep 17, 2026
Policy· 2 min read

New York Orders Utilities to Disclose All AI Use Within 60 Days

State regulators cite risks from hallucinations, bias, and cybersecurity vulnerabilities as utilities deploy AI for customer service and maintenance.

Via AI Watch · Sep 17, 2026
Policy· 3 min read

AI Liability Without Regulation Is Insufficient, Marcus Argues

A growing push from tech investors to embrace liability while rejecting oversight ignores lessons from aviation, pharmaceuticals, and finance.

Via AI Watch · Sep 17, 2026