TechNewsReel
Live

AI Data Hunger Drives Bulk Acquisition of Rare, Out-of-Print Books

Independent booksellers report massive orders for obscure titles suspected of being scraped for LLM training.

TechNewsReel Newsroom · August 12, 2026

Independent bookstores across the globe are reporting a surge in unusually large orders for obscure, out-of-print titles, sparking fears that AI companies are harvesting physical texts to train Large Language Models. These acquisitions target rare materials that have not yet been digitized, effectively turning physical archives into proprietary training data.

Reports of these bulk purchases have emerged from bookstores in Germany, Spain, the United States, New Zealand, and Australia, with some orders exceeding a thousand volumes. In Galway, Ireland, the bookstore Kennys received what operators described as a "bananas" order for 5,000 obscure titles. The bookstore's operators believe these volumes are destined for AI programmers seeking unique datasets. A Canadian firm, Zoom Books, has been specifically identified as a primary actor in these worldwide acquisitions of second-hand and out-of-print materials for scraping purposes.

The Hunt for Dark Data

This trend is driven by the industry's pursuit of "dark data"—physical texts that remain offline or out of print. As AI developers face an increasing wave of copyright lawsuits and legal challenges over the use of digital data, the incentive to acquire physical copies has grown. By purchasing physical books, companies can potentially bypass digital copyright hurdles or secure specialized knowledge that is not available in existing web-scraped datasets.

Risks to Bibliographic Heritage

The shift toward bulk physical acquisitions raises significant concerns regarding the preservation of global bibliographic heritage. Industry observers warn that if rare and out-of-print books are purchased in bulk solely for digitization, the physical history of literature and specialized regional knowledge could be permanently lost. While some reports suggest a pattern of "destructive scanning" where books are destroyed after digitization, Zoom Books has denied these allegations, characterizing its operations as a standard recycling and trading model. Regardless of the fate of the physical copies, the transition of these texts from public or archival access to proprietary AI datasets represents a fundamental shift in how knowledge is stored and accessed.

Future Outlook

As the demand for high-quality, non-redundant training data increases, the pressure on independent booksellers and rare book archives is expected to grow. The industry is now watching whether this trend will trigger new protective measures for physical archives or lead to further legal disputes over the ownership of digitized content derived from physical acquisitions. For now, the scale of the orders in places like Galway serves as a signal of the lengths to which AI companies will go to feed their models.

Sources

Get a notification when a big story breaks. A few a day at most — no spam.