虎嗅

After computing power, comes language data: Is the new bottleneck in AI reshaping the pricing power of the content industry?

原文:算力之后是语料:AI的新瓶颈正在重塑内容产业的定价权?

Summary of the Core Content

This news article focuses on the central contradiction of a shortage of high-quality training data for large AI models: As the scale of AI model parameters expands (from hundreds of billions to trillions) and models evolve towards multi-modal capabilities (with a surge in the demand for images, videos, and interactive data), the internet's native high-quality data is gradually depleting. Additionally, AI-generated content can contaminate the data pool, posing a risk of "model collapse." As a result, capital has flocked to the A-share market sectors related to AI training data (such as Readker Culture and CITIC Publishing). However, traditional content providers (publishers) face multiple barriers in transitioning from selling books to selling data, including complex copyright issues, an absence of pricing mechanisms, and the potential for AI to replace their business models. The ultimate conclusion is that the scarcity of training data is stratified—professional, exclusive, and multi-modal data are truly in short supply. Publishing companies have yet to realize significant revenue from AI, and their long-term competitiveness will depend on their ability to continuously create high-quality content and monetize their data assets.

Detailed Analysis

1. Why Are Large AI Models Facing a "Fuel Shortage"? – The Crisis of High-Quality Training Data

Large AI models are like data refineries, where computing power represents the furnace and data the raw material. However, the quality of the data is more critical than the amount of computing power available. Unlike computing power, which follows Moore's Law (doubling every 18 months), data becomes increasingly scarce:

  • Exponential Demand: GPT-2 required dozens of GB of text for training, while GPT-3 needed hundreds of GB; modern multi-modal models (capable of understanding images, text, and robot interactions) require even more data. For instance, embodied intelligence models need at least 10 million hours of multi-modal data, but in China, only 500,000 hours of compliant physical interaction data are available, leaving a 99% gap.
  • Depleting Sources: Free human-generated content on the internet (such as blogs, forums, and books) is being rapidly depleted. Moreover, AI-generated content (e.g., articles written by AI) can contaminate the training datasets.

In short, AI consumes data at a much faster rate than humans can produce it.

2. The "Boomerang Effect" of AI-Generated Content

What happens when AI-generated content is used to train further AI models? It's like copying copies of copies—each generation loses some information, leading to model degradation (for example, the model may fail to understand rare events or produce incorrect results). This phenomenon is known as "model collapse."

  • Mitigation Strategies: Not all AI-generated content is problematic. By strictly filtering out AI-generated parts, mixing it with human-generated data, and cleaning the dataset, the degradation can be mitigated.
  • Silicon Valley's Efforts: To obtain clean training data, Anthropic secretly scanned books worldwide (under the "Panama Project"). They were sued for downloading 7 million books from pirated sources and settled for $1.5 billion, highlighting the high value and difficulty of obtaining clean, original data.

3. Traditional Publishers Turning into "AI Miners" – The Transition from Selling Books to Selling Data

Publishers used to earn money by selling books, but now AI companies need high-quality data. They can license copyrighted content for use in model training. However, there are several key considerations:

  • Dramatic Price Differences: The price of data varies significantly depending on its purpose (e.g., training models vs. retrieval). For example, HarperCollins and Microsoft licensed old books for $5,000 per book every three years, with the revenue split equally between the publisher and the author, solely for training purposes.
  • Value of Content: Ordinary novels and publicly available ancient texts are plentiful and less valuable, whereas specialized, exclusive, or industry-specific content (e.g., medical textbooks, legal cases, industry reports) are in high demand and therefore more profitable.
  • Current Situation in China: AI companies are purchasing data from publishers, but publishers must first digitize, clean, and annotate the data, and also handle author authorization issues (many authors of older books are unreachable, breaking the authorization chain).

4. Alternative Paths for AI Companies

AI companies are not waiting for publishers to provide data; they have alternative strategies:

  • Synthetic Data: They can generate high-quality data on their own (e.g., mathematical problems, code), which can replace some real data in certain applications.
  • Improving Data Efficiency: They are developing more efficient model architectures (e.g., MoE) and training methods to use less data for better results.
  • RAG (Retrieval-Augmentation) Techniques: Models don't need to memorize all knowledge during pre-training; they can retrieve information from databases on demand (e.g., asking an AI about the GDP in 2026).
  • Other Sources: They purchase data from specialized service providers, use publicly available data, or directly contact authors to obtain custom data.

Therefore, traditional book copyrights are just one of many potential data sources for AI companies.

5. From Copyright to Pricing Power – Three Barriers Hindering Publishers

Even if publishers hold copyright, they may not have control over pricing due to three major obstacles:

  • Unclear Copyright Rights: Publishers only hold the right to publish; the actual copyright often belongs to the authors. Licensing data for AI training requires individual author agreements, which is difficult to obtain, especially for older books.
  • Lack of Pricing Mechanisms: There are no standardized rules for calculating the amount of data used by AI models or for determining how much to pay authors. Currently, most overseas transactions involve purchasing entire datasets, not per word.
  • Alternative Data Sources: The availability of synthetic data, specialized services, and open-source data reduces publishers' bargaining power. Additionally, the free distribution of some books (e.g., ancient texts) lowers the price of regular publications.

Conclusion: To gain pricing power, publishers must continuously produce high-quality original content and transform their copyrights into compliant, traceable, and tradable data assets. They also need to establish stable profit-sharing mechanisms with AI companies and authors. Relying solely on existing text collections is not enough to maintain a competitive advantage in the AI era.

Additional Note: The Surge in A-Share Market Prices

The sudden rise in the AI training data sector in August was not due to publishers actually making profits from AI but rather on speculation about the scarcity of training data. Currently, the systems for rights confirmation, pricing, and profit-sharing are not well-established, and most publishers' revenue from AI is still at the agreement stage. In the long run, the demand for multi-modal data (images, videos, robot interactions) will increase, while the value of pure text data will decrease. The true scarcity lies in specialized, exclusive, and compliant multi-modal data, as well as ongoing human creativity. If top creators cannot benefit from AI licensing agreements, the motivation to produce high-quality content will decline, potentially affecting the entire industry's development. Therefore, a healthy balance between AI companies, publishers, and authors is essential for the industry's sustainable growth.