TechNewsReel
Live

AI Labs Buy and Scan Physical Books to Escape 'Slop' Contamination

Companies are purchasing pre-2022 printed books through brokers like ISBNdb, scanning them with destructive methods, and discarding the originals to secure uncontaminated training data.

TechNewsReel Newsroom · July 27, 2026

AI companies are systematically purchasing physical books printed before 2022, scanning their contents for training data, and destroying the originals—often by slicing off spines to feed pages through high-speed scanners.

The practice represents an escalating response to "model collapse," the degradation that occurs when AI systems train on their own synthetic output. As the internet floods with AI-generated text, labs are hunting for "slop-free" human-authored content that predates the LLM era.

The optics problem

ISBNdb, which facilitates bulk book purchases for AI companies, openly acknowledges the reputational risks. "The world's best AI training data is sitting on a shelf," the company states. Internal marketing copy notes that "'AI company destroys two million books' is not a headline that generates sympathy."

To shield clients from backlash, ISBNdb offers strict non-disclosure agreements and buyer anonymity. The broker facilitates bulk orders ranging from hundreds to potentially millions of volumes per transaction, according to reporting from The Next Web and 404 Media.

Why pre-2022 books?

Books printed before 2022 are specifically targeted because they predate widespread AI-generated content and modern data poisoning techniques. These volumes offer high-density, edited, and authoritative text that cannot be contaminated by synthetic data circulating online.

The destructive scanning process—slicing spines to process pages rapidly—is chosen for speed and cost efficiency over non-destructive alternatives. Once digitized, the physical books are discarded.

Legal landscape

The practice received legal validation when Judge William Alsup ruled in Bartz v. Anthropic that training AI on lawfully purchased books qualifies as fair use, describing it as "quintessentially transformative." The ruling cited first-sale doctrine principles regarding disposal rights of legally acquired materials.

However, the legal picture remains complicated. Anthropic separately paid a $1.5 billion settlement over pirated books from Library Genesis, a distinct issue from the legal purchasing program. Court documents referenced in secondary reporting mention an internal initiative called "Project Panama" focused on book acquisition, though Anthropic has not publicly confirmed this designation.

In 2024, Anthropic hired Tom Turvey, a former Google Books executive, to lead book acquisition operations—signaling continued investment in physical text procurement.

Cultural cost

The practice raises questions about the preservation of physical literary history versus the pursuit of cleaner training data. Rare and antique books are being permanently destroyed for marginal improvements in model performance, transforming cultural artifacts into disposable data sources.

As one industry observer noted, the tension between digital preservation and aggressive data acquisition highlights how AI's hunger for training material now manifests in physical destruction—not just digital scraping.

Sources

Get a notification when a big story breaks. A few a day at most — no spam.