AI Firms Scan and Destroy Post-1970 Books to Avoid Training on AI Text
Companies are buying books by ISBN, digitizing them destructively, then claiming copyright exemptions—but the 'rare book' framing may be overblown.
AI training drives mass book scanning
AI companies have launched large-scale operations to scan books published before 2002, seeking training data guaranteed to be free of AI-generated content. The process is destructive: books are purchased, unbound for easier digitization, scanned, and then destroyed. The practice has drawn criticism, particularly around claims that "rare books" are being lost, according to a detailed analysis first reported by Hackaday.
The scanning targets books identifiable by ISBN numbers—the international cataloging system introduced only in 1970. This detail undermines the narrative that priceless historical manuscripts are at risk. These are mass-produced modern books, not medieval illuminated texts.
The companies appear to be pursuing a copyright strategy: by buying a physical book, digitizing it, and destroying the original, they hope to argue that only one copy exists, potentially sidestepping publisher copyright claims. The approach mirrors failed legal strategies from companies like ReDigi and Aereo in previous decades.
Why it matters
The controversy reveals two significant trends. First, AI companies are already contending with "model collapse"—the degradation that occurs when AI systems train on AI-generated content. Their retreat to pre-2002 books suggests the internet's utility as a training corpus has been compromised by the proliferation of AI-generated text. Second, the debate exposes how easily symbolic concerns about book destruction can overshadow substantive questions about AI companies' data practices, copyright circumvention, and competitive hoarding of digitized knowledge.
The economics of book destruction
The publishing industry routinely pulps vast quantities of books. Unsold inventory, outdated editions, and low-value backlist titles are recycled by the ton annually. Most books printed since 1970 exist in sufficient quantities that their physical scarcity doesn't equate to cultural loss.
A book can be technically uncommon without being valuable. Small print runs of genealogical works or niche technical manuals may have few surviving copies, but if they've been considered disposable for decades, their sudden destruction for AI training doesn't represent a new threat to cultural heritage.
The real concern isn't the physical destruction—it's what happens to the digital copies. These scans remain in private corporate archives, inaccessible to researchers, libraries, or the public. Unlike digitization efforts by institutions like the Internet Archive, this knowledge extraction serves only proprietary model training.
Competitive data hoarding
Multiple AI companies are conducting simultaneous scanning operations, potentially competing to digitize the same titles. This raises questions about whether firms are deliberately purchasing multiple copies of legitimately scarce books to deny competitors access—a practice that would align with the winner-take-all dynamics of the AI industry.
The absence of any commitment to share these scans with public archives like Archive.org means the cultural benefit of digitization accrues entirely to private shareholders.
Jenny List, writing for Hackaday with experience in the publishing industry, argues the "rare book" framing distracts from more substantive criticisms of AI companies' practices. The symbolic power of book destruction, she notes, has inflated the perceived value of what are largely common, post-1970 publications with ISBN numbers—books that would otherwise face pulping in the normal course of publishing economics.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call
