AI Training on Copyrighted Books Hinges on Fair Use, Not Copying
Recent court rulings reveal judges are treating AI model training more like reading than reproduction, but the legal landscape remains unsettled.

AI Training on Copyrighted Books Hinges on Fair Use, Not Copying
The question of whether AI companies can legally train their models on copyrighted books has no simple answer, despite the intuitive sense that using authors' works without permission should be illegal. Recent court decisions reveal a complex legal landscape where the distinction between reading and copying may determine the future of the AI industry.
Why it matters
These early rulings are shaping how billions of dollars in AI investment will flow and whether generative AI tools can continue operating as they do today. For business leaders evaluating AI adoption, the unsettled legal environment creates risk around both using AI tools and building AI-powered products that rely on training data.
A Landmark Ruling That Favored AI Companies
Last year, Judge William Alsup ordered Anthropic to pay $1.5 billion to writers whose works trained the company's AI models—but not because the training itself was illegal. According to TechCrunch, which first reported these details, Alsup ruled that Anthropic's AI training was lawful. The penalty was for obtaining books from illegal shadow libraries, not for using them to train models.
"Like any reader aspiring to be a writer, Anthropic's LLMs trained upon works not to race ahead and replicate or supplant them — but to turn a hard corner and create something different," Judge Alsup wrote, comparing AI training to how writers study literature.
Cathy Gellis, an attorney specializing in intellectual property and technology, told TechCrunch the ruling benefits AI companies. "Copyright law hinges on copying, but it doesn't hinge on using the work or experiencing the work, consuming the work, reading the work," she explained.
The Fair Use Question
Courts are applying fair use doctrine—a carve-out in copyright law allowing limited use of copyrighted material for criticism, parody, education, and other purposes—to determine whether AI training is permissible. The analysis turns on whether the use is "transformative" enough and whether it competes with the original work's market.
Jason Henderson, Senior Attorney and Founder of the IP & Media Practice at JWL International, told TechCrunch that courts are finding patterns. "If what you're doing is you're training on somebody's property because your purpose is to directly compete, then the courts will frown on it," he said. "If what you're doing is not going to compete, then the courts are tending to find ways that it will be okay."
This principle emerged clearly in Thomson Reuters v. Ross Intelligence, where Judge Stephanos Bibas ruled last year that training on Reuters' content to build a competing legal platform was not fair use because it lacked a transformative purpose.
The Challenge of Outdated Law
Judges are interpreting copyright guidelines last updated in 1976 to address technology questions that didn't exist then. "The law is all over the place, and it's because of this question," Henderson said. "They know that the AI model has been trained on so much stuff, and the law has not really caught up to that question."
Most AI companies remain in pending litigation, meaning definitive answers won't arrive soon. "It would be kind of foolish for the AI companies to ignore them," Gellis said of the early rulings, even as later decisions could reverse their influence.
These details were first reported by TechCrunch.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call

