虎嗅

AI companies are snapping up old books from before 2022: The more machines can write, the more valuable human data becomes.

原文:AI公司抢购2022年前的旧书:机器越会写,人类数据越值钱

Summary of Key Points

AI companies are now competing to purchase old books published before 2022. The reason is that the amount of content generated by AI on the internet (referred to as “AI waste”) is increasing, and using this material to train models can lead to what’s known as a “model collapse” – similar to copying a document over and over again, resulting in increasingly blurred details. Old books, being written by humans and unable to be altered by AI, represent authentic and reliable training data. There is also a legal loophole: purchasing physical books, scanning them into digital copies, and then destroying the original books is considered “fair use,” which is safer than downloading pirated versions. intermediaries like ISBNdb saw this as a business opportunity and helped AI companies find old books in bulk, but later discontinued their service (possibly due to public relations considerations). This incident highlights a new trend in the AI era: data that is genuine, has a proven source, and can be verified as not being machine-generated is the most valuable.

1. Why Do AI Companies Suddenly Prefer Old Books? Because of the Surplus of “AI Waste” on the Internet

You might wonder: hasn’t AI already consumed all the information on the internet? Why do they still need to search for old books?

The answer lies in the year 2022 – ChatGPT was released in November 2022, and since then, a large amount of content generated by AI (such as articles, product descriptions, and paper abstracts) has begun to circulate on the internet. The content you see online today may be several times removed from its original source, similar to “waste” that lacks value.

If AI uses this material to train new models, what will happen? It’s like copying a document repeatedly: the first copy might still be clear, but by the third iteration, details start to disappear, and eventually, it becomes unrecognizable. This is what’s meant by a “model collapse” – AI becomes less efficient and loses its ability to process complex information and creativity.

Old books published before 2022, on the other hand, are different. They were written by humans and printed on paper, making them immune to AI manipulation. Even if they’re not of high quality, they represent “pure human-generated” data.

2. How Does the Legal Loophole Work? Buying Physical Books for Scanning and Destruction Is Legal

Here’s a case that goes against common sense: Claude’s parent company, Anthropic, once used pirated e-books to train models. Fearing legal issues, they came up with a clever strategy:

They spent millions of dollars on physical books, removed the spines, scanned them quickly into digital copies, and then destroyed the original books, keeping only the digital versions (this was called the “Panama Project”).

An American judge ruled that this approach was legal! The judge’s reasoning was as follows:

1. Using books to train models is considered fair use;

2. Purchasing physical books, converting them into digital copies, and destroying the originals without distributing them further is also fair use;

3. However, directly downloading pirated PDFs and storing them in a database is illegal.

In simple terms, as long as you buy the books with real money and create only one internal digital copy that isn’t shared, it’s not considered copyright infringement. Anthropic’s strategy not only avoided the risk of piracy but also obtained high-quality data – quite clever, indeed.

3. Intermediaries See the Business Opportunity: Finding Old Books for AI Companies Has Become a New Industry

Where there is demand, there is a market. ISBNdb, a well-established company that previously provided book information (such as ISBN numbers and inventory), now specializes in helping AI companies find old books:

  • They search for books from used bookstores, libraries, and out-of-print catalogs based on the AI companies’ lists;
  • They can purchase anywhere from 1,000 to 1 million books at a time;
  • They also act as intermediaries to keep the publishers unaware that the purchases are from AI companies.

However, after reports by 404 Media, they removed their service page, claiming it was just an exploration of market demand. But this is likely just a public relations statement; as long as AI companies continue to need clean data, this business will likely continue.

4. Genuine Data Is the True Treasure: Only What’s Not Machine-Generated Is Valuable

The most important aspect of this phenomenon is not how many books are destroyed by AI, but rather what it reveals about the scarcity in the AI era: genuine, verifiable data.

Old books are valuable not because they contain a large amount of information, but because they go through a “human selection process” – someone selects the topic, writes the content, edits it, and buys it (readers make the final decision with their money). This process isn’t perfect, but it leaves behind a record of real-world choices that are not fabricated by machines.

This principle applies to real data in everyday life as well:

  • Product reviews on clothing e-commerce sites (real user feedback, not AI-generated);
  • Production failure data from factories (actual errors that AI cannot create);
  • Medical case records (real medical conditions, not simulations).

In the future, the most valuable data will not be secret or massive in volume; it will be the data that can clearly explain its origin, who reviewed it, and why it’s credible.

In Conclusion

As AI becomes more capable of generating content, data that is not machine-generated becomes increasingly precious. The story of old books serves as a reminder: in an era where machines can produce endless text, human creativity and choices are the most irreplaceable values.