Summary of Key Points
American AI companies, such as Amazon and Anthropic, are purchasing large quantities of second-hand and old books (including rare and out-of-print titles) to address the shortage of training data for large-scale models. After scanning the content, they destroy the physical books. Behind this practice is the AI's urgent need for "high-quality, authentic human data." Additionally, due to a U.S. court ruling that allows the retention of electronic copies as part of "fair use" after the destruction of physical books, this approach has become both compliant and cost-effective. However, it has also sparked controversy over the loss of rare knowledge and the public's access to it.
1. Why Do AI Companies Compete for Old Books?
AI models require a vast amount of "authentic content written by humans," but the internet data has been largely depleted:
- Exhausted Old Data: Content from old newspapers, forums, and blogs has been extensively used by AI companies.
- Contaminated New Data: Much of the online content is generated by AI, and using such data to train new models can lead to "model collapse"—similar to humans becoming increasingly foolish if they continuously consume their own waste.
- **Old Books as "Golden Data": Books from decades ago lack AI-generated content and have undergone multiple rounds of review by authors and editors, making them of high quality. Many niche and specialized books have not been digitized, representing fresh sources of data for AI.
Thus, old books, which were once neglected, have suddenly become highly sought after by AI companies.
2. Why Destroy the Books After Scanning?
AI companies are not philanthropes; their decision to destroy books is well-thought-out:
- Legal Compliance: A U.S. court ruling has determined that AI companies can legally purchase a physical book, scan it, and destroy the original, treating it as if the electronic copy replaced the physical one (without infringement).
- Efficiency: Removing the spine allows for faster scanning of hundreds of pages at once, which is crucial when dealing with millions of books.
- Cost Savings: Destroying the books eliminates the need for storage and transportation costs.
For example, Anthropic's "Panama Project" involved processing 500,000 to 2 million books in six months, with destruction being the most cost-effective option.
3. The Destruction of Rare Books: A Crisis of Knowledge Loss
The most concerning aspect is the destruction of rare and out-of-print books:
- Irrenewable Loss: These books may be unique in existence, and their destruction means they are lost forever.
- Knowledge Privatization: Once purchased and scanned by AI companies, the content becomes part of their internal databases, inaccessible to the public. Digitalization should have made knowledge more accessible, but instead, it has turned public knowledge into private assets for AI companies.
Even Elon Musk has raised concerns, urging the xAI team to use non-destructive scanning methods (like Google Books) to preserve the original books in libraries. However, these methods are less efficient and more costly, so AI companies may be reluctant to adopt them.
4. Legal Rulings: Encouraging "Book Destruction" Indirectly?
U.S. court rulings have played a pivotal role in this situation:
- No Path for Piracy: Anthropic faced legal issues when downloading electronic books from pirated websites and was fined $1.5 billion.
- **Destruction as the "Legitimate Option": The courts have approved the practice of purchasing and destroying physical books while retaining electronic copies, providing AI companies with a compliant alternative.
This has created a paradox: stealing books (piracy) is costly, while purchasing and destroying them is now legal. The law does not explicitly prohibit book destruction, but it has inadvertently given it a green light.
5. Future Concerns: More "Book Destruction"?
If more AI companies follow this approach:
- Standard Practice: Destroying books may become a common practice to avoid copyright risks and reduce costs.
- Increased Knowledge Privatization: More human knowledge from the pre-AI era could become trapped within AI companies' closed datasets.
- Disappearance of Rare Knowledge: Out-of-print books that have not been digitized could be permanently lost.
Who knows whether Musk, who currently criticizes Anthropic, will also opt for destruction in the future due to the pursuit of efficiency? After all, developing better AI models is the ultimate goal for these companies.
At heart, this issue reflects a conflict between the "needs of AI development" and the "protection of human knowledge." AI requires data to thrive, but we must prevent precious knowledge from being lost in the process. Balancing these two aspects will be a critical issue in future AI regulation.