虎嗅

Why are Anthropic and Amazon destroying physical books?

原文:Anthropic和亚马逊为什么要销毁正版图书?

Summary of Key Points

Some American AI companies (such as Anthropic and Amazon) are purchasing large quantities of second-hand paper books to avoid copyright risks. These books are scanned into digital text for model training, after which the original copies are destroyed. This practice stems from specific provisions in U.S. copyright law, whereas Chinese AI companies do not need to follow this approach due to legal differences. Additionally, this “book destruction” may also serve competitive purposes aimed at monopolizing exclusive data and could potentially lead to the loss of physical carriers of human civilization.

Detailed Analysis

1. How do American AI companies “destroy books”?

The process is akin to an assembly line operation in a “book dismantling factory.” Anthropic’s book destruction project, called “Project Panama,” involves buying large numbers of second-hand, genuine books from the market, transporting them to a processing center, cutting off the spines, separating the pages, and scanning each page into digital text for AI training. The original paper books are then disposed of. Amazon has also adopted this method, with second-hand book sellers receiving an increasing number of bulk purchase orders.

However, this is not technically necessary; for example, SpaceX’s AI uses careful scanning techniques to avoid damaging the spines of valuable books, and Google’s early Google Books project did not destroy the scanned books. In essence, “book destruction” is a strategic decision made by these companies rather than a last resort.

2. Why does U.S. copyright law force companies to destroy books?

U.S. law allows you to buy a genuine book and resell it, give it away, or even destroy it (the “first sale rule”). However, scanning the book into a digital format while keeping the physical copy could be considered “illegal copying” of the author’s rights. By destroying the physical book after scanning, companies can argue that they have merely converted the format from paper to digital without creating an additional copy, thus reducing copyright risks.

Anthropic previously faced a huge lawsuit for training its model using 7 million books from pirated websites, resulting in a $1.5 billion settlement. Therefore, it is more cost-effective for them to spend millions on purchasing and destroying second-hand books to avoid such legal issues.

3. Are there other motives beyond compliance?

Some suspect that companies like OpenAI and Google may destroy books to monopolize exclusive content. If rare or out-of-print books are destroyed, only their AI models would contain that information, giving them a competitive advantage. After all, more unique training data can potentially enhance the performance of AI models.

4. Why don’t Chinese AI companies destroy books?

Chinese copyright law is different from U.S. law. Regardless of whether the scanned books are destroyed or not, companies must obtain permission from the authors or publishers to use the digital content for AI training (since scanning itself constitutes a form of copying). Therefore, even if Chinese companies tried to follow the American practice, they would still face copyright violations. Destroying the books would be unnecessary and wasteful.

5. The concern about book destruction: Could it lead to the loss of human civilization?

Paper books are physical carriers of human civilization. If AI companies destroy them, and those companies go out of business or their models stop being used, the content of those books could be lost forever. For example, a scanned and destroyed rare book would represent a permanent loss of cultural heritage. This “destructive” approach to data acquisition essentially consumes the physical backups of human civilization.

In summary, American AI companies destroy books due to legal requirements and business considerations, but this practice carries risks related to the preservation of cultural heritage. Chinese companies do not need to follow this path because of different legal frameworks. This incident also highlights that the development of AI should not come at the expense of human cultural heritage.