Unsealed Filings Reveal Anthropic's Secret 'Project Panama' Book Scanning Operation
Court documents disclose a covert effort by the AI firm to acquire and destructively scan millions of physical books for training data.
Anthropic executed a secret operation known as "Project Panama" to acquire millions of physical books for the purpose of training its artificial intelligence models. The details of the initiative emerged from unsealed court filings in a copyright lawsuit, revealing a systematic effort to digitize physical texts while keeping the operation hidden from public view.
According to the filings, the company purchased vast quantities of books from used-book sellers. Once acquired, the books underwent a process of "destructive scanning," where every page was digitized for AI training data and the physical remains were subsequently recycled. Internal documents describe the project as an ambitious effort to "destructively scan all the books in the world." The level of secrecy surrounding the operation was explicit; internal planning materials stated, "We don't want it to be known that we are working on this."
The Race for High-Quality Data
This operation comes amid a high-stakes competition among AI developers to secure high-quality training data. While much of the internet has already been scraped, books are considered a superior source of information for improving a model's reasoning and writing capabilities. By targeting physical copies, AI companies can access a wealth of structured, long-form knowledge that may not be available in a clean, digital format or may be protected by digital access controls.
Legal and Ethical Implications
Project Panama reveals a physical-world acquisition strategy designed to bypass traditional digital copyright hurdles. By purchasing physical copies of books, the company attempted to create a legal buffer between the acquisition of the text and the digital reproduction of the content. However, the practice of destructive scanning—destroying the original work to create a digital copy—raises significant legal and ethical questions regarding the fair use of copyrighted materials and the transparency of how frontier AI models are built.
Future Outlook
As the copyright lawsuit progresses, the industry is watching whether this "physical-to-digital" pipeline will be deemed a permissible method of data acquisition or a systemic violation of intellectual property rights. The revelation of Project Panama highlights a growing tension between the insatiable data needs of large language models and the legal protections afforded to authors and publishers. It remains to be seen if other AI labs have employed similar clandestine physical acquisition strategies to gain a competitive edge.