虎嗅

Old books are starting to gain value again.

原文:旧书又开始值钱了

Summary of Key Points

The training of large AI models has depleted the internet's linguistic data, leading to a cycle where AI-generated content feeds back into the models, resulting in overfitting (the generated content becoming monotonous and lacking authenticity). In this context, old books published before 2022 have suddenly become highly sought after. These books have not been contaminated by AI, have undergone a professional publishing process, and contain obscure and rare information that can address the data challenges faced by AI training. The value logic of the used book market has shifted: once neglected items are now targeted by AI companies willing to pay high prices for them, signaling an unexpected resurgence for the publishing industry.

Detailed Analysis

1. Why Do AI Models Suddenly Crave Old Books?

The reason is that internet data has become heavily contaminated:

  • Circular Pollution: With the rise of generative AI, much of the content on the web is created by AI, which leads to a homogenization of the material (e.g., everyone in AI-generated short videos has the same expressions, and articles use repetitive sentence structures like “not... but...”).
  • Severe Overfitting: Hugo Award winner Hao Jingfang noted that half of novels are written by AI, yet authors dislike them because the resulting works lack originality and aesthetic value.

Old books published before 2022 have largely remained untouched by AI, providing uncorrupted data for model training.

2. Are Old Books More Reliable Than Online Content?

The publishing process provides a valuable pre-processing service for AI:

  • Human Review: Authors write long manuscripts, editors select content, and publishers ensure accuracy, creating a structured and high-quality dataset.
  • Better Information Quality: Old books focus on a single theme with consistent terminology and coherent arguments, in contrast to online pages filled with ads, broken links, and SEO spam.

3. How Have Obscure Old Books Become Valuable?

Traditional used book markets often feature unsold titles, but AI training has changed this:

  • Bestsellers Are No Longer Useful: Classics like *Dream of the Red Chamber* and *One Hundred Years of Solitude* have been digitized multiple times, making them less valuable for model training.
  • Rare Knowledge: Obscure books, such as local railway manuals or 90s software guides, contain unique terms and facts that AI models need to avoid redundancy.

A 2024 study in *Nature* confirmed that preserving rare data can prevent model degradation. For example, Spanish book sellers found buyers willing to pay for Catalan-language old books despite the high shipping cost.

4. Who Is Buying Old Books?

The new players in the used book market are AI companies:

  • AI Companies: They spend millions on purchasing and scanning paper books into PDF format. Although they were fined for selling pirated e-books, legal paper copies are recognized by courts.
  • Strange Orders: Dutch book sellers have received requests for specific languages and versions at prices higher than the book value; these likely reflect AI companies accumulating rare data.

5. Where Could This Trend Spread?

This phenomenon could extend to other undigitized resources:

  • Old Newspapers/Magazines: Contain unique language and historical information.
  • Handwritten Materials: Diaries, traditional craft records, etc.
  • Minority Languages: Old books in minority languages or dialect recordings.
  • Physical Archives: Corporate reports, government documents.

In essence, AI seeks “rare information produced by humans that has not been altered by AI.” Any resource meeting this criteria could become a target for AI companies.

Conclusion

AI’s demand for clean data has given the struggling publishing industry a new lease on life through old books. This also highlights the potential value of forgotten, undigitized resources.