Amazon Destroys Rare Books After Scanning for AI Training Data
A 404 Media investigation tracked a rare volume to a Nevada warehouse where workers slice off covers and discard originals after digitization.

Amazon's book-destruction pipeline revealed
Amazon operates a warehouse in Las Vegas where workers process rare and out-of-print books solely to create AI training data, destroying the physical volumes in the process. The operation came to light through a 404 Media investigation that placed a tracking device inside a rare book and followed it to the company's VGT3 facility in Nevada.
Employees at the warehouse told investigators their job consists of receiving bulk book deliveries, slicing off covers to speed scanning, and running pages through digitization equipment. Once scanned, the books become unusable. The team working on this project has adopted a logo featuring a dinosaur holding a book.
Amazon confirmed the practice in a statement, saying the company "purchase[s] books through commercial channels to improve the products and services customers use," according to TechCrunch.
Why it matters
The destruction of rare physical books for digital training data highlights the collision between preserving cultural artifacts and feeding AI systems' insatiable appetite for content. As AI companies exhaust readily available internet text, they're turning to increasingly aggressive acquisition strategies — including purchasing materials that exist in limited quantities and may have historical or scholarly value. This raises questions about whether short-term AI development goals justify permanently removing rare texts from circulation.
The strategic value of pre-digital content
Rare and out-of-print books serve a specific purpose in AI training pipelines. These texts often exist nowhere on the internet, providing unique linguistic patterns and content unavailable through web scraping. They also offer a critical quality guarantee: anything published before 2022 predates widespread generative AI availability, meaning it couldn't have been produced by a language model.
This distinction matters because feeding AI-generated text back into training datasets can trigger what researchers call model collapse — a progressive degradation in output quality as models learn from their own synthetic productions rather than authentic human writing.
Broader data acquisition strategies
The book-destruction operation represents one front in a wider campaign by AI companies to secure training material. Amazon has also moved to harvest content from its streaming platform Twitch, automatically enrolling all creators in a program that allows the company to use streams, clips, images, and chat logs for AI training unless users manually opt out.
Competitors are pursuing parallel strategies. Google is paying $10 million to acquire internal business data from bankrupt Spirit Airlines, a trove that includes roughly 100 million emails and 500 million Microsoft Teams chats. Some publishers have negotiated formal licensing arrangements instead: Amazon's agreement with the New York Times for AI-focused content is valued between $20 million and $25 million annually.
The book-buying and destruction operation had not been publicly reported before 404 Media's investigation.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call
