AI Companies Buying Books in Bulk to Scan and Destroy
Independent bookstores report unusual bulk orders as 404 Media tracks a book's journey to an AI training operation.
Investigation uncovers AI training book pipeline
Artificial intelligence companies are purchasing books in unexpectedly large quantities from used and independent bookstores, scanning them to train AI models, and then destroying the physical copies, according to a new investigation by 404 Media.
The investigative outlet took the unusual step of placing a tracking device inside a book to document its journey from bookstore shelf to its final destination. Emanuel Maiberg, a journalist and co-founder at 404 Media, discussed the findings in an interview with CBS News.
Independent booksellers have reported receiving bulk orders that struck them as unusual in both size and pattern. The investigation suggests these purchases are part of a systematic effort to acquire training data for large language models and other AI systems.
Why it matters
The practice raises questions about copyright, fair use, and the economics of AI training data. While companies have scraped publicly available text from the internet, purchasing and destroying physical books represents a different approach to data acquisition. Independent bookstores may unknowingly be supplying material that helps build commercial AI systems without authors' knowledge or compensation. The investigation also highlights the lengths AI companies are willing to go to secure training data as they compete to build more capable models.
Tracking books to their destination
By embedding tracking technology in a book, 404 Media was able to follow the physical path of a purchased volume. The tracking revealed the book's movement from the point of sale through the supply chain to what investigators believe was a scanning operation.
The method provides concrete evidence of what had previously been suspected based on unusual purchasing patterns reported by bookstore owners. These booksellers had noticed orders that didn't fit typical consumer or reseller behavior.
The broader context of AI training data
AI companies require massive amounts of text to train language models. While much of this data has come from web scraping, books represent high-quality, long-form content that can improve model performance. The practice of buying physical books to digitize them sidesteps some of the technical challenges of acquiring digital copies, though it raises its own legal and ethical questions.
The destruction of books after scanning suggests the primary value lies in the digital text rather than resale of physical copies, indicating these operations are focused specifically on data extraction for AI purposes.
The investigation was first reported by 404 Media, with findings discussed in a CBS News interview with co-founder Emanuel Maiberg.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call
