Policy

AI Companies Are Buying Used Books to Train Language Models

A copyright settlement has cleared the way for tech firms to purchase physical books and destroy them for training data.

Omega Editorial· September 3, 2026· 2 min read

AI firms turn to physical books for training data

Artificial intelligence companies have begun purchasing used books on a large scale to feed their language model training pipelines, according to Fast Company. The practice has accelerated following a copyright settlement earlier in 2026 that provided legal clarity for the controversial approach.

The settlement established that AI companies can legally buy physical books and destroy them in the process of digitizing their contents for model training purposes. Courts ruled this practice falls under fair use provisions of copyright law, despite concerns from authors and publishers about the wholesale destruction of books for commercial AI development.

The trend has created noticeable effects in the used book market, with volumes disappearing from circulation across multiple countries as AI firms compete for training material.

Why it matters

This legal precedent fundamentally reshapes how AI companies can access copyrighted text for training data. While digital piracy and unauthorized scraping have drawn lawsuits, the physical book loophole offers a legally defensible path—albeit one that permanently removes books from circulation. The ruling may accelerate the arms race for quality training data while raising questions about cultural preservation and authors' ability to control how their work is used in AI systems.

Legal framework enables new data strategy

The court settlement represents a significant shift in how copyright law applies to AI training. By purchasing physical copies, companies establish clear ownership of the material before digitization, distinguishing this approach from scraping online content or using library databases without permission.

For AI developers, physical books offer curated, edited text that has passed through traditional publishing quality controls—potentially more valuable for training than raw internet data. The practice also provides a paper trail demonstrating legitimate acquisition, important as copyright litigation continues across the industry.

Market implications

The scale of book purchasing by AI companies remains unclear, but the impact on used book availability suggests substantial investment in this data acquisition strategy. As large language models require massive text corpora for training, the demand for physical books could continue growing, particularly for works not readily available in digital formats.

The details were first reported by Fast Company.

#artificial intelligence#large language models#copyright law#training data#fair use#book market

This is an original analysis by the Omega editorial team. Source reporting: AI Watch.

Want systems like this working for your business?

Book a Call

More in Policy

Policy· 3 min read

NYC and LA Ban Student AI Use While SF Expands AI Tutors

The nation's largest school districts impose restrictions as San Francisco delays policy decisions and scales up AI literacy programs.

Via AI Watch · Sep 4, 2026
Policy· 3 min read

New York lawmakers push Hochul to fast-track AI workforce bills

State legislators argue that waiting until year-end to act on AI regulation carries greater risk as the technology evolves daily.

Via AI Watch · Sep 4, 2026
Policy· 3 min read

Wisconsin Law Enforcement to Enforce State AI Child Abuse Law

Local officials say they will continue prosecuting AI-generated child sexual abuse material cases despite federal appeals court ruling.

Via AI Watch · Sep 3, 2026