Microsoft Exec Called AI Scraping 'Largest Theft of Labor'
Newly unsealed court documents reveal internal admissions that AI training practices constitute theft and pose existential threats to publishers.

Internal Documents Contradict Public Defense
Newly unredacted court filings in The New York Times' copyright lawsuit against OpenAI and Microsoft have exposed internal communications that directly undermine the companies' public legal defense. The documents, unsealed as part of the three-year-old case, show executives privately acknowledging that their AI training practices amounted to theft while publicly arguing for fair use protections.
According to the filings, Microsoft's Director of Applied Science, Brent Hecht, described the practice in a January 2023 internal memo as "an astonishing theft of unprecedented proportions" and "the largest theft of labor in human history." The documents were first reported by TechCrunch.
Scale and Methods of Content Acquisition
The unsealed materials reveal the scope of copying for the first time. OpenAI's training datasets contain more than 91,692 copies of works from The New York Times, Daily News, and Center for Investigative Reporting. A single dataset derived from Common Crawl included over 2 million documents from nytimes.com alone.
The filings detail specific methods allegedly used to acquire content. OpenAI researcher Nick Ryder reportedly told company president Greg Brockman about a "hack to get around nytimes paywall," to which Brockman replied "ah nice." The companies also allegedly stripped copyright notices from training data before feeding it to models, with researchers noting they "wouldn't want model outputting" copyright notices to users.
Microsoft provided training data to OpenAI through initiatives called Project Taxi and Project Mango, with the latter containing copies of at least 160,903 unique works from news publishers.
Admissions of Market Harm
The documents include statements that directly challenge fair use defenses, which require that new uses not substitute for or harm the market of original works. Microsoft's own data shows its Copilot answer engine caused click-through rates for The New York Times to drop as much as 93% compared to traditional Bing search.
An internal Microsoft presentation described this decline as a "doom loop" that would "hurt the performance of our models and the entire web at the same time." The presentation noted it is "highly unusual that an end-product threatens the economic foundations of its essential suppliers."
OpenAI's Head of ChatGPT, Nick Turley, wrote internally that publishers face an "existential threat" from chatbots, which are "largely substitutive" and "will get more and more substitutive as they get better." Microsoft CEO Satya Nadella testified under oath that he would have required OpenAI to retrain its models had he known they scraped paywalled content.
Why It Matters
These internal admissions could significantly weaken the fair use defense that has so far protected AI companies in copyright litigation. While judges have generally sided with AI firms, evidence that executives privately acknowledged both the theft-like nature of their practices and the direct market harm to content creators may shift legal outcomes. The case could establish precedent for how AI companies must compensate or license content from publishers, potentially reshaping the economics of both industries.
The details were first reported by TechCrunch based on newly unredacted portions of court filings. Neither OpenAI nor Microsoft responded to requests for comment.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call