虎嗅

$1.5 billion settlement in Anthropic book copyright dispute marks the beginning of an AI-driven pricing era for the publishing industry: A comprehensive review of the largest AI-related copyright settlement in U.S. history and its implications for China's publishing sector.

原文:十五亿美元Anthropic图书版权和解案终裁,出版业迎来AI定价时代:美国史上最大AI版权和解案全程复盘及其对中国出版业的启示

Summary of Key Points

The largest AI copyright settlement in U.S. history, the "Bates v. Anthropic" case, has come to an final decision: The AI company Anthropic was ordered to pay $1.5 billion (to be paid in four installments) for downloading nearly 500,000 books from pirated websites to train its models. Each rights holder received approximately $3,000 per book. This settlement is not only the largest amount ever obtained by the publishing industry from the AI sector but also sets a global benchmark for determining the value of content in the AI era, providing a roadmap for future AI copyright litigation. It also serves as a wake-up call for China's publishing industry to establish a legal infrastructure for AI content transactions as soon as possible.

Detailed Explanation

1. The Key Figures of the $1.5 Billion Settlement: $3,000 per Book, with an Acceptance Rate of Over 90%

The figures in this case are astonishing:

  • $1.5 billion in compensation: To be paid in four installments; the first two amounts of $600 million have already been received, with interest added if overdue (9 percentage points higher than the national debt rate), and Anthropic cannot get back a single cent even if no one claims the rights.
  • $3,000 per book: After deducting attorney fees, each rights holder receives around $3,000, which is four times the amount in typical infringement cases and 15 times that in cases of unintentional infringement (significantly higher than the settlement reached with Google Books back then).
  • 92.77% acceptance rate: While the acceptance rate in class-action lawsuits is usually around 10%, this case reached nearly 93%, with only 54 opponents (half of whom wanted to be included in the claim list), indicating that the settlement offer was very attractive.
  • Attorney fees cut in half: The lawyers initially demanded $300 million (20%), but the judge awarded only $101 million (6.8%), a record for copyright litigation yet still reasonable.

The significance of these figures lies in the fact that this is not just an arbitrary payment; it represents the compliance cost that AI companies must bear for using pirated data and also provides the industry with an understanding of the value of data used for AI model training.

2. The Turning Point of the Case: The Judge's Distinction Between "Training" and "Acquisition"

The victory in this case was largely due to the judge's distinction between two aspects:

  • Legal training: The court ruled that using legal books to train AI models constitutes "transformative use" (e.g., converting the book content into knowledge for the model), which is permitted by law as fair use.
  • Pirated acquisition: Anthropic's act of downloading and storing books from LibGen (a pirated website) was considered a direct infringement, regardless of the purpose of training, and thus subject to higher legal penalties.

Why did Anthropic agree to pay $1.5 billion? If it had lost the case, the statutory compensation would have been $150,000 per book for 500,000 books, totaling $75 billion, which would have led to bankruptcy. Therefore, the settlement was a way to minimize losses.

3. A New Approach for Future AI Copyright Cases: Targeting the Source of Piracy Instead of Focusing on Fair Use

This case sets a precedent for other plaintiffs:

  • The "shadow library" strategy: Instead of arguing about the legality of training, plaintiffs can directly investigate the source of data used by AI companies (whether it comes from pirated websites). For example, in Meta's case, the court recognized the legality of training but still upheld the charge of piracy.
  • Higher standards for fair use: Even if training is legal, if an AI company has access to legally licensed content and chooses not to use it, it will be harder to argue for fair use. Litigation is driving AI companies to purchase legitimate content.
  • From confrontation to negotiation: News organizations and record labels have already signed licensing agreements with OpenAI and Meta. The $1.5 billion settlement amount serves as a reference for future negotiations, showing AI companies the potential costs of not purchasing licensed content and content owners the potential gains from litigation.

In short, AI companies will now be more cautious about using pirated data; they must either purchase legitimate licenses or face legal action.

4. Implications for the Publishing Industry and Authors

  • For the publishing industry:
  • Short-term benefit: Leading publishers like HarperCollins could earn tens of millions of dollars, a significant boost to their profits.
  • Long-term impact: The $3,000 per book price becomes a reference point for AI model training, and future negotiations will likely be based on this figure. More importantly, metadata (such as ISBNs and copyright registrations) become valuable tools for generating revenue, as only books with clear ownership can be monetized.
  • For authors:
  • $3,000 per book represents a relatively certain payout, although lower than the statutory maximum of $150,000. Settlements provide a safer option compared to the risks associated with litigation.
  • This case highlights the importance of copyright registration: Unregistered books may not even be eligible for compensation. Organizing as authors (e.g., through associations) can lead to better terms in settlements.
  • Authors should also retain their rights for future use of their work, as AI-generated derivative works may still be infringing.

5. What China’s Publishing Industry Should Learn

China’s publishing industry should not wait for litigation to act but should establish a legal infrastructure for AI content:

  • No hope for a "Chinese version of the Bates case": China does not have the same statutory compensation levels or class-action capabilities as the U.S., so relying on litigation for large payouts is unrealistic; licensing agreements are necessary.
  • Digitize rights information: Publishers need to document copyright details, authors' contact information, and contract terms for each book to establish clear ownership.
  • Revised contracts: Traditional contracts often lack clarity regarding AI training and data usage rights; licenses should specify these terms clearly (e.g., dividing profits 50-50 as in the U.S. case).
  • Export considerations: If books are intended for the U.S. market, copyright registration is essential to participate in settlements.
  • Collaboration: Individual publishers need to join forces through associations to negotiate with AI companies and create premium products using licensed Chinese content.

In summary, this case marks a shift from allowing unrestricted use of pirated data to establishing a legal framework for AI and content. For China’s publishing industry, it is crucial to improve rights management to prevent high-quality Chinese content from being freely copied or sold at low prices. The battle over AI and content usage is just beginning.