虎嗅

**Anthropus settles for $1.5 billion: The case is over, but the legal issues are just beginning.**

原文:Anthropic15亿美元和解:案件已了,法律问题才刚开始

Summary of Key Points

Anthropic paid a settlement of $1.5 billion to resolve copyright issues related to the use of pirated books to train its AI model Claude, covering approximately 500,000 works at a cost of around $3,000 per work. This is considered the largest copyright settlement in the AI era. However, the crucial point is that this amount represents a payment for the use of pirated data, not for legitimate training authorization. The settlement does not set a binding legal precedent, and authors have only received limited compensation. Meanwhile, the U.S. Congress is pushing legislation to require AI companies to disclose their training data, leaving the future rules for AI training copyright unresolved.

1. $1.5 Billion is a Copyright Fine, Not Money for AI Training

To understand this settlement, it’s important to consider a 2025 court ruling: using legitimate data (such as scanned copies of books purchased legally) for AI training is considered fair use and does not require payment. In contrast, downloading and storing over 7 million books from pirated repositories like LibGen constitutes an infringement.

Why would Anthropic pay $1.5 billion? According to U.S. copyright law, the maximum penalty for intentional infringement is $150,000 per work, meaning a potential compensation of $1 trillion for 7 million works. The $1.5 billion was seen as a way to mitigate the risk of exorbitant damages at an affordable cost.

Therefore, the $3,000 per work does not represent authorization fees for training AI but rather a settlement for using pirated materials. If Anthropic had used legitimate data, it might not have had to pay anything to the authors at all.

2. This Settlement is Not a Legal Final Decision; Other Cases May Still Be Decided Differently

Anthropic claims to have won because the judge ruled that using legal data for training is fair use. However, in U.S. law, only a final judgment from an appeals court constitutes a binding precedent. This settlement was reached during the pretrial phase, and the judge’s opinion is merely advisory.

There are still dozens of AI copyright cases pending: for example, the New York Court is reviewing Google Gemini, and the California Court is reviewing Meta’s practices. Different judges have varying views on whether AI training infringes on authors’ rights. Additionally, six authors withdrew from the settlement to file separate lawsuits, and 350 major publishers/right holders opted to pursue their own claims, believing they could obtain higher compensation. This indicates that no party is fully satisfied with the outcome.

3. Authors Have Not Earned Much; Money Is Highly Deducted

Although $1.5 billion seems substantial, authors receive very little:

  • Costs Deducted: $122 million is set aside for attorney fees, litigation costs, and management expenses (which could amount to 25%-30% based on initial estimates).
  • Distribution with Publishers: The remaining money is split equally between authors and publishers.
  • Smaller Authors Suffer More: With larger publishers withdrawing from the settlement, the majority of remaining authors receive only around $1,500 per work, which is less than 2% of the statutory maximum compensation.

Many authors are dissatisfied, arguing that the compensation is too low or the distribution is unfair. This hardly qualifies as a victory; they have essentially received a fraction of what they are entitled to after being infringed upon.

4. Congress Aims to Increase Transparency in AI, Leading to More Challenges

The U.S. Congress is proposing two bills aimed at increasing transparency around AI training data:

  • CLEAR Act: Before releasing products, AI companies must submit a list of copyrighted works used in their training data to the copyright office, which will then make this information public.
  • TRAIN Act: Copyright holders can request subpoenas to verify whether their works have been used in AI training.

These bills do not directly determine whether AI training is infringement but grant copyright holders the right to investigate. In the future, holders will easily identify who has used their works and can either sue or negotiate licensing fees with AI companies, making it harder for them to use pirated materials covertly.

5. The Case Is Over, But the AI Copyright War Has Just Begun

This settlement is merely a stopgap measure for Anthropic; the battle over AI training copyright is far from over:

  • The judge’s ruling on fair use is not binding, and future cases may still determine that AI training constitutes infringement.
  • The transparency measures proposed by Congress will expose AI companies to more lawsuits.
  • The $3,000 per work settlement may not set a standard; future compensation or licensing fees could be higher.

For large AI companies, the real challenge lies in determining whether and how much of their training data can be made public. If each work requires separate licensing agreements, the costs could become prohibitively high.

In summary, the $1.5 billion settlement is just a minor setback; the struggle over AI and copyright rights is just entering its most intense phase.