Microsoft Staff Called AI Training 'Largest Theft' in Unsealed Docs
Internal communications reveal deep concerns about AI models scraping copyrighted content as copyright lawsuit proceeds.

Internal warnings about AI training practices
Microsoft employees raised alarm bells about OpenAI's use of copyrighted news articles, with one internal document describing it as potentially "the largest theft of labor in human history," according to court filings unsealed Thursday in the ongoing copyright lawsuit brought by The New York Times and other publishers.
The documents, reported by The New York Times, emerged from the case filed in late 2023 that now includes twelve publishers suing both Microsoft and OpenAI. Judge Sidney Stein is currently reviewing summary judgment motions in the Southern District of New York.
One 2023 Microsoft memo warned that millions worldwide would soon view large AI models "hoovering up" their work as "an astonishing theft of unprecedented proportions." The same author noted that these models "are a product that destroys its supply chain."
Microsoft identified the author as Brent Hecht, a director of applied science who also held a Northwestern University position. The company stated these views don't represent official positions and that Hecht was employed specifically to "present divergent and asymmetric perspectives" rather than make decisions.
Executive testimony and internal communications
Microsoft CEO Satya Nadella testified that "anything that is paywalled should be licensed by anyone who wants to use it." He stated he would have exercised Microsoft's contractual right to force OpenAI to retrain its models had he known about training on paywalled content. A company spokesman clarified Nadella "spoke to broad principles" about information access.
At OpenAI, internal messages showed staff discussing workarounds for publisher paywalls. When one employee told president Greg Brockman about building a "hack" to bypass The New York Times paywall, Brockman responded: "ah nice."
Nick Turley, who led the ChatGPT team, wrote in June 2023 that AI posed an "existential threat" to publishers. By February 2024, he noted AI products "will get more and more substitutive as they get better," later stating that AI "products are largely substitutive, period."
An OpenAI engineer observed in February 2023 that "no matter how prominently we show the links, users won't click"—undermining arguments that chatbots drive traffic back to original publishers.
Why it matters
These unsealed documents provide rare insight into what technology companies privately understood about the impact of their AI training practices, even as they publicly defended them. The internal acknowledgments that AI models could substitute for original content and harm publishers directly contradict fair use arguments both companies are making in court. For business leaders considering AI deployment, the case highlights growing legal and ethical scrutiny around training data sources.
Legal arguments and broader implications
Both Microsoft and OpenAI maintain their training constitutes fair use, arguing they transform articles into new work rather than substitute for originals. The case has already compelled OpenAI to preserve 20 million ChatGPT conversation logs.
Steven Lieberman, representing the New York Daily News and seven other publications, said: "The world can see what OpenAI and Microsoft thought all along about the fairness of their own behavior."
A 2020 memo from then-policy director Jack Clark warned leadership the company was "creating systems that substitute for the labor of the people that define the 'culture' of society" and would "become the symbol of how Silicon Valley is thoughtlessly stepping into other parts of life and leaving a mess on the carpet." Clark later left to co-found Anthropic.
These details were first reported by The New York Times based on newly unsealed court documents.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call