AI Companies Bulk-Buy Pre-2022 Books to Avoid AI-Generated Text
A database firm now helps labs acquire millions of printed volumes guaranteed free of the synthetic content flooding the web.

AI labs turn to printed books for clean training data
Artificial intelligence companies are purchasing old printed books by the thousands—and sometimes millions—because they offer something increasingly rare on the internet: content guaranteed to be free of AI-generated text.
ISBNdb, which operates what it describes as the world's largest book database, now offers bulk book acquisition services specifically targeting AI companies. The firm markets printed books published before 2022 as ideal training material, noting they are "structurally guaranteed" to be free of the synthetic text now saturating online sources.
The pitch reflects a growing challenge for AI developers: as large language models flood the web with generated content, finding clean training data becomes harder. Training AI systems on AI-generated text can trigger "model collapse," where models become less accurate and more error-prone with each generation.
How the book-buying operation works
ISBNdb facilitates orders ranging from 1,000 to one million books per transaction. The company's metadata makes it easier for AI labs to systematically acquire, scan, and digitize printed volumes while avoiding duplicate purchases.
The service emphasizes discretion. ISBNdb advertises that it maintains strict non-disclosure agreements and never reveals client identities or acquisition strategies. The company's website acknowledges the sensitivity directly: "'AI company destroys two million books' is not a headline that generates sympathy."
Booksellers on marketplaces including Alibris and Biblio have reported unprecedented bulk purchases starting in April. One professional bookseller specializing in foreign language titles told 404 Media he went from selling roughly 20 books weekly to hundreds per week. The orders showed no thematic pattern beyond all books having ISBNs, and buyers showed "total disregard for the price of the book"—characteristics suggesting AI company purchases rather than traditional library acquisitions.
Legal questions and destructive scanning
The practice gained public attention in January when a copyright lawsuit against Anthropic revealed internal documents detailing plans to acquire and scan millions of printed books. Court filings showed Anthropic contracted with Datamation, which offers both destructive and non-destructive scanning. Destructive scanning cuts book spines to feed pages through machines—faster and cheaper than preserving the physical volume.
Federal Judge William Alsup ruled Anthropic's digitization legal specifically because the physical books were destroyed afterward. He found creating one digital copy to replace a purchased print original constituted fair use, since the digital version wasn't shared outside the company.
ISBNdb markets this legal framework to potential clients, arguing that purchasing used books "does not deprive any creator of income they would otherwise have received" and represents "the completion of a book's lifecycle."
Why it matters
The shift to printed books as training data reveals how quickly AI-generated content has contaminated the web—the very platforms these models were designed to learn from. It also raises preservation concerns: rare and foreign language books destroyed during scanning may become permanently harder to obtain. The practice creates a paradox where AI companies must reach backward to pre-digital sources to avoid the pollution their own technology creates, while potentially eliminating physical copies of uncommon works in the process.
These details were first reported by 404 Media.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call
