AI Firms Buying Secondhand Books in Bulk for Training Data
Independent booksellers report mysterious warehouse shipments as court ruling enables destructive scanning for language models.

Mysterious bulk orders flood secondhand book market
Independent booksellers across multiple countries are reporting an unprecedented surge in bulk book purchases, with orders shipping to distant warehouses rather than individual readers. Stuart Manley of Barter Books in Northumberland recently received a single order from a Canadian company equivalent to his typical weekly sales of two to three thousand books—a pattern he hasn't witnessed in three decades of trading.
The unusual purchases span diverse genres with no apparent logic, ranging from obscure Latin texts to cowboy novels. Booksellers suspect these volumes aren't destined for human readers but for training artificial intelligence systems.
Court ruling opens door to book destruction
The surge follows a 2025 US court decision that fundamentally changed how AI companies can source training material. Judge William Alsup ruled that using purchased books to train AI software constitutes fair use under copyright law, calling the practice "exceedingly transformative."
The case involved AI firm Anthropic and three authors who sued over the use of their work. Unsealed court documents revealed that Anthropic's training process for its Claude chatbot involves destroying books through what the company internally called "Project Panama"—an initiative aimed at "destructively scanning all the books in the world," according to the filings first reported by the BBC.
Destructive scanning removes book spines to enable rapid page-by-page digitization at industrial scale, with the physical remains then recycled. An Anthropic spokesperson confirmed the company trains Claude on "publicly available web data, commercially acquired datasets, and data we generate ourselves," noting that sourcing books for training is standard across the AI industry. The company stated it does not purchase rare or antiquarian books for destruction.
Why it matters
This development creates a collision between AI's hunger for diverse training data and cultural preservation concerns. While copyright law in the US now permits this practice, UK law requires copyright owner permission for such copying, according to Oxford University intellectual property specialist Professor Emily Hudson. The legal divergence could fragment how AI companies source training material globally, and raises questions about whether unique or historically significant editions might be lost to scanning operations despite company assurances.
Booksellers face ethical dilemma
The situation presents a complex trade-off for sellers. David Tobin of Walden Books in north London appreciates moving inventory that has sat unsold for years, but expresses concern about ultimate destruction. Derek Walker of Edinburgh's McNaughtan's bookshop distinguishes between recent academic texts with library copies and genuinely rare 18th-century editions where only one example survives.
Manley acknowledges the recycling benefit for unwanted titles, noting "the world no longer needs five million copies of The Da Vinci Code." He reports selling books advertised online for two decades that found no buyers until now.
Experts suggest the random subject matter supports the AI training theory—unusual and rare texts provide fresh material that could improve large language model performance beyond commonly available digital content.
These details were first reported by the BBC.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call
