Policy

Amazon Warehouse Scans and Destroys Rare Books for AI Training

Investigation reveals industrial book-destruction operation at Nevada facility, raising questions about copyright and cultural preservation.

Omega Editorial· August 18, 2026· 3 min read

Amazon's Book-Scanning Operation Exposed

Amazon is purchasing and destroying rare books on an industrial scale to train artificial intelligence models, according to an investigation by 404 Media. The operation came to light when a bookseller grew suspicious of a 1,000-book order placed through Biblio, a marketplace that allows anonymous purchases. To trace the shipment's destination, the seller planted an Apple AirTag inside one volume.

The tracker led to a massive Amazon warehouse in Nevada designated VGT3, where employees confirmed their primary task is scanning books—a process that requires removing spines and destroying the originals. "All we do is scan books," one Amazon worker wrote on an employee forum. "Some are assigned to cut books, and others go to receive where they get books and scan the barcodes."

Amazon provided only a brief statement to 404 Media: "Amazon purchases books through commercial channels to help develop and improve the products and services our customers use." The company did not address questions about the destruction of rare materials.

Why it matters

This practice sits at the intersection of three critical issues: the preservation of physical cultural artifacts, the boundaries of copyright law in the AI era, and the transparency of tech companies' data acquisition methods. A recent court ruling found that Anthropic's similar book-scanning operation was "transformative" and therefore legal—a precedent that could accelerate the destruction of rare texts for AI training data. For businesses investing in AI, this reveals both the aggressive data-gathering strategies of major players and the unresolved legal questions surrounding training data sources.

Legal Precedent and Industry Practice

The book-scanning practice became public knowledge through a lawsuit against Anthropic, which revealed the AI company was using industrial equipment to cut pages from millions of books, scan them, and dispose of the originals. A judge ruled this process didn't violate copyright law because converting physical texts to digital format qualified as "transformative."

This legal framework has apparently emboldened other AI companies. Booksellers told 404 Media they've long suspected certain orders were destined for AI training, citing telltale signs: unusually large orders, seemingly random book selections, and the requirement that all volumes have ISBNs (International Standard Book Numbers).

Scanning Every Book Ever Printed

The ISBN detail has fueled a concerning theory among booksellers: AI companies may be attempting to scan every printed book in existence by systematically working through ISBN databases. The fact that Amazon warehouse workers scan ISBNs of all processed books lends credence to this hypothesis, according to 404 Media's reporting.

The bookseller who participated in the investigation noted that the destroyed books were rare—meaning few copies remain in circulation. "There are different types of value," the seller explained. "There's monetary value, obviously, but there are a lot of other types of value. There's historical value, intellectual value, sentimental value. All sorts of things, and all of those the AI companies don't care about. They just want the content as a bunch of words strung together."

The warehouse logo itself seems to acknowledge the operation's nature: a T-Rex holding a book it's about to consume—an image that, ironically, appears AI-generated.

These details were first reported by 404 Media, which conducted the AirTag investigation and interviewed Amazon warehouse employees and booksellers.

#amazon#ai training data#copyright#book preservation#anthropic#data acquisition

This is an original analysis by the Omega editorial team. Source reporting: AI Watch.

Want systems like this working for your business?

Book a Call

More in Policy

Policy· 4 min read

Pentagon AI Targeting Tempo Outpaces Evaluation Capacity

The Defense Department is deploying frontier AI models on 30-day timelines while the infrastructure to independently test them remains years behind.

Via AI Watch · Aug 18, 2026
Policy· 3 min read

AI-Written NIH Grants Win More Funding But Favor Safe Ideas

New analysis of 125,000 federal research proposals reveals generative AI increases success rates while potentially narrowing scientific exploration.

Via AI Watch · Aug 18, 2026
Policy· 4 min read

AI Infrastructure Boom Fueled by Off-Balance-Sheet Debt

Special purpose vehicles and complex financing structures are spreading AI buildout risk across the financial system, raising questions about stability.

Via AI Watch · Aug 18, 2026