Amazon is destroying rare books to feed its AI training models
The e-commerce giant is reportedly pulping out-of-print texts after scanning them to avoid AI-generated 'slop'.
Amazon is buying massive quantities of rare and out-of-print books, scanning them for AI training data, and destroying the physical copies in the process. This operation highlights a desperate race among AI labs to secure high-quality human text as the internet's available data is exhausted.
An investigation by 404 Media tracked a shipment of rare books to an Amazon facility in Las Vegas, identified as VGT3, which is dedicated to this scanning and destruction operation. According to reports from TechCrunch and Ars Technica, the process involves cutting off the bindings of the books to allow for rapid page scanning, after which the original physical copies are pulped or shredded.
The Hunt for 'Clean' Data
AI companies are specifically targeting rare and out-of-print titles because they provide long-form, high-quality training data that predates the 2022 surge of large language models. This ensures the data is free of AI-generated "slop," or contaminated content produced by other LLMs, which can degrade the performance of new models. While some firms claim to collaborate with institutions to preserve texts, others have opted for destructive scanning to maximize speed and reduce operational costs.
Cultural Erasure for Corporate Gain
This practice represents a significant moral and cultural conflict, as irreplaceable historical artifacts are permanently lost to advance corporate AI capabilities. The trend has sent shockwaves through the rare book trade, where sellers are seeing a surge in bulk orders for obscure titles. Tomás Kenny of Kennys Bookshop described the impact of these developments, stating that AI is "terrifying our industry."
A Pattern of Secrecy
Amazon is not the only firm accused of these tactics. Reports from Ars Technica indicate that Anthropic utilized a codename, "Project Panama," to conceal its own destructive book-scanning operations. These efforts underscore the extreme lengths to which AI labs will go to solve the "data wall" problem—the point at which models run out of new, high-quality human-written text to learn from.
What's Next
As AI firms continue to scour the physical world for data, the industry faces increasing scrutiny over the ethics of data acquisition. It remains to be seen if regulators or cultural heritage organizations will intervene to protect remaining physical archives from being converted into digital training sets and then destroyed.