Amazon Destroys Rare Books to Feed Frontier AI Models, Investigation Finds
A 404 Media probe reveals the e-commerce giant is stripping bindings from rare volumes at a Las Vegas facility to accelerate data ingestion.
Amazon is systematically destroying rare books to harvest training data for its frontier artificial intelligence models, according to an investigation by 404 Media. The operation involves the physical destruction of rare and foreign-language volumes to facilitate high-speed digitization.
The investigation uncovered that Amazon is utilizing a facility in Las Vegas, Nevada, designated as VGT3, to process these materials. To accelerate the scanning process, employees at the warehouse cut off the bindings of the books, effectively destroying the physical integrity of the volumes. 404 Media confirmed the destination of these books by placing a tracking device—specifically an AirTag—inside a shipment of rare books and following it directly to the Amazon warehouse. The team operating the VGT3 facility reportedly uses a logo featuring a T. rex brandishing its teeth while holding or devouring a book.
A Pattern of Industrial Digitization
This industrial approach to data acquisition stands in stark contrast to the methods used by organizations like the Internet Archive, which digitizes books page-by-page to preserve the physical object. However, Amazon's tactics echo the early 2000s, when Google Books employed similar destructive scanning methods to build its massive digital library. The current surge in AI development has reignited this tension, as companies race to acquire massive datasets to improve the reasoning and linguistic capabilities of their frontier models.
The Cost of 'Content'
The practice highlights a systemic conflict where the immediate needs of AI development are prioritized over the preservation of cultural and intellectual heritage. By treating rare books as mere "content"—strings of words to be ingested—AI companies ignore the historical, sentimental, and intellectual value inherent in the physical medium.
"There are different types of value... monetary value, obviously, but there are a lot of other types of value," an unnamed bookseller noted, explaining that historical and sentimental values are disregarded by AI firms that only seek raw data. This shift represents a permanent loss of artifacts that cannot be recovered once the bindings are severed and the volumes are processed.
Legal and Ethical Fallout
Beyond the physical loss, the operation raises significant legal questions regarding intellectual property. The use of rare books for AI training is often defended under the umbrella of "fair use," yet the scale of this operation suggests a massive ingestion of copyrighted material without compensation or consent.
As AI companies continue to exhaust available public data, the industry is watching to see if other tech giants adopt similar destructive methods. It remains to be seen whether regulators or copyright holders will successfully challenge these practices in court, or if the physical destruction of rare libraries will become a standard cost of the AI arms race.