Policy

AI Companies Buying Books in Bulk to Scan and Destroy

Independent bookstores report unusual bulk orders as 404 Media tracks a book's journey to an AI training operation.

Omega Editorial· August 19, 2026· 3 min read

Investigation uncovers AI training book pipeline

Artificial intelligence companies are purchasing books in unexpectedly large quantities from used and independent bookstores, scanning them to train AI models, and then destroying the physical copies, according to a new investigation by 404 Media.

The investigative outlet took the unusual step of placing a tracking device inside a book to document its journey from bookstore shelf to its final destination. Emanuel Maiberg, a journalist and co-founder at 404 Media, discussed the findings in an interview with CBS News.

Independent booksellers have reported receiving bulk orders that struck them as unusual in both size and pattern. The investigation suggests these purchases are part of a systematic effort to acquire training data for large language models and other AI systems.

Why it matters

The practice raises questions about copyright, fair use, and the economics of AI training data. While companies have scraped publicly available text from the internet, purchasing and destroying physical books represents a different approach to data acquisition. Independent bookstores may unknowingly be supplying material that helps build commercial AI systems without authors' knowledge or compensation. The investigation also highlights the lengths AI companies are willing to go to secure training data as they compete to build more capable models.

Tracking books to their destination

By embedding tracking technology in a book, 404 Media was able to follow the physical path of a purchased volume. The tracking revealed the book's movement from the point of sale through the supply chain to what investigators believe was a scanning operation.

The method provides concrete evidence of what had previously been suspected based on unusual purchasing patterns reported by bookstore owners. These booksellers had noticed orders that didn't fit typical consumer or reseller behavior.

The broader context of AI training data

AI companies require massive amounts of text to train language models. While much of this data has come from web scraping, books represent high-quality, long-form content that can improve model performance. The practice of buying physical books to digitize them sidesteps some of the technical challenges of acquiring digital copies, though it raises its own legal and ethical questions.

The destruction of books after scanning suggests the primary value lies in the digital text rather than resale of physical copies, indicating these operations are focused specifically on data extraction for AI purposes.

The investigation was first reported by 404 Media, with findings discussed in a CBS News interview with co-founder Emanuel Maiberg.

#ai training data#copyright#independent bookstores#large language models#investigative journalism#404 media

This is an original analysis by the Omega editorial team. Source reporting: AI Watch.

Want systems like this working for your business?

Book a Call

More in Policy

Policy· 2 min read

Judge to rule on Musk company challenge to Minnesota AI law

A hearing Tuesday addressed efforts to block the state's ban on AI-generated nonconsensual intimate imagery.

Via AI Watch · Aug 19, 2026
Policy· 3 min read

AI Adoption Surges, Yet U.S. Job Market Remains Stable

Despite rapid AI deployment across industries, unemployment hasn't risen—raising questions about what comes next for workers.

Via AI Watch · Aug 19, 2026
Policy· 3 min read

Musk's xAI Challenges Minnesota Ban on AI Nudification Tools

Company argues state law restricting deepfake nude image generators is overly broad and violates free speech protections.

Via AI Watch · Aug 19, 2026