Policy

AI Firms Scan and Destroy Rare Books to Train Language Models

An Apple AirTag investigation traced antique books from sellers to Amazon warehouses where they're digitized and shredded.

Omega Editorial· August 17, 2026· 3 min read

Major artificial intelligence companies are purchasing rare and antique books by the thousands, scanning them to train large language models, then destroying the physical copies—a practice that has ignited controversy among authors, archivists, and preservationists.

Court documents from a lawsuit against Anthropic revealed the company's "Project Panama," an initiative to buy millions of physical books, remove their bindings, scan the pages at high speed, and shred the originals. Internal planning documents stated the goal plainly: "destructively scan all the books in the world," according to reporting first published by Forbes.

Industrial-scale book destruction

An investigation by 404 Media used an Apple AirTag to track nearly 1,000 books from a rare book seller through multiple facilities before arriving at an Amazon-owned warehouse in Colorado Springs. The building features a Tyrannosaurus rex logo preparing to eat an open book, and employees describe their work as simply scanning books all day.

The scale is substantial. Unsealed court records show Anthropic planned to scan between 500,000 and 2 million books under a single six-month vendor contract. Anonymous sources told 404 Media that bulk orders range from 1,000 to 1 million books per purchase, with used booksellers across the United States, United Kingdom, and Europe reporting routine orders of hundreds to thousands of titles.

Why physical books matter for AI training

AI developers have shifted to physical books to solve two critical problems with digital training data. First, using pirated online archives exposes companies to copyright litigation from authors and publishers. Purchasing physical books legally allows the owner to use that copy as they choose, including destroying it.

Second, physical books offer higher-quality training material. Companies specifically target books printed before 2022 to ensure the text is human-authored rather than AI-generated. The internet now contains substantial AI-written content, and training models on that material doesn't improve their capabilities. Physical books provide edited, long-form human reasoning and narrative structure that short web pages lack.

The practice also grants access to millions of out-of-print, specialized, and niche books that exist only in physical form.

Why it matters

While the First Sale Doctrine makes this practice legal, it creates a preservation dilemma. Rare and antique books are being permanently destroyed to create proprietary training datasets, with no guarantee the digitized versions will be preserved or made accessible. For books that exist in limited physical copies, this represents an irreversible loss to the historical record—all to feed models that may be obsolete within years. The tension between legal rights and cultural preservation will likely intensify as AI companies exhaust readily available training data.

Legal but controversial

The practice remains lawful under copyright law. The First Sale Doctrine allows book owners to destroy their purchases, and courts have ruled that converting lawfully purchased books into digital files for model training constitutes "transformative" fair use. Critics argue companies are exploiting physical ownership as a loophole to build datasets without compensating creators, but no legal prohibition currently exists.

Anthropic's internal documents noted the company didn't want it known they were pursuing this strategy, suggesting awareness of potential public backlash.

These details were first reported by Forbes and 404 Media, which conducted the AirTag tracking investigation.

#artificial intelligence#large language models#copyright law#book preservation#anthropic#training data

This is an original analysis by the Omega editorial team. Source reporting: AI Watch.

Want systems like this working for your business?

Book a Call

More in Policy

Policy· 3 min read

Most AI Companies Miss California Detection Tool Deadline

An independent audit reveals that five of 13 major generative AI platforms have not published required content verification systems.

Via AI Watch · Aug 17, 2026
Policy· 3 min read

Illinois Man Charged With Using AI to Create Child Abuse Images

Kane County prosecutors allege Jeremy Batterman generated and altered explicit material of minors using artificial intelligence tools over a six-month period.

Via AI Watch · Aug 17, 2026
Policy· 3 min read

SEC Clears Path for Nvidia's $500B AI Data Center Financing Push

New regulatory guidance exempts certain data center debt from post-crisis risk retention rules, potentially reshaping infrastructure investment.

Via AI Watch · Aug 17, 2026