虎嗅

OpenAI accused of hiding crucial evidence: Does the fair use defense in the New York Times case hang in the balance?

原文:OpenAI被指隐匿关键证据,《纽约时报》案的合理使用抗辩悬了?

Summary of Key Points

In its copyright lawsuit against OpenAI, The New York Times has accused OpenAI of concealing crucial user conversation logs, misleading the court (by claiming it could not search for copyrighted material in its training data), and even deleting billions of records. The Times is calling for severe sanctions against OpenAI. These logs are central to OpenAI’s defense, which argues that the use of copyrighted material in its ChatGPT training constitutes fair use. If the concealment of evidence is confirmed, OpenAI could face not only legal penalties but also a collapse of its core argument. Both Chinese and U.S. laws impose strict penalties for concealing or destroying evidence in lawsuits. Google has previously lost a monopoly lawsuit due to similar behavior, and Chinese law stipulates fines, detention, and even criminal charges for such actions.

Why Does The New York Times Want to Sanction OpenAI?

In short, OpenAI tried to outsmart the process with its evidence, but was caught in the act:

  • Manipulated Log Submission: It was agreed that OpenAI would provide 20 million anonymized user conversation logs to verify whether ChatGPT used copyrighted material from The New York Times and whether it reproduced the original text. However, the logs provided by OpenAI were heavily edited and unusable.
  • Lies Exposed: OpenAI had previously told the court that it was too costly and a violation of privacy to search its training data for NYT content, but its own engineers revealed that the company had actually searched for this information before the lawsuit.
  • Evidence Destruction: The Times alleges that OpenAI deleted billions of conversation records after the lawsuit, violating the court’s order to preserve evidence.

The Times is demanding that the court impose strict sanctions: either OpenAI must admit that these logs are useless or the court will assume that the concealed/deleted content is detrimental to OpenAI (which would effectively mean admitting the truth).

Why Are These Logs So Critical for OpenAI?

OpenAI’s main defense is based on the concept of fair use: “We are training ChatGPT to understand language patterns, not to copy original texts; any reproduction of original text is a technical glitch.” However, these logs could undermine this argument:

1. The Amount of NYT Content in the Training Data: If the logs show that the training set contains a large number of complete NYT articles, the court may question whether OpenAI had no choice but to use this material and whether it purchased the rights legally, potentially threatening its business.

2. Frequency of Reproductions: OpenAI claims that such reproductions are rare, but if the logs prove that ChatGPT frequently and accurately reproduces NYT articles, it would indicate that OpenAI has not only learned the language but also stored the original texts. In a commercial context, this would constitute direct infringement rather than fair use.

3. Risks with the RAG (Retrieval, Aggregation, and Generation) Model: ChatGPT’s RAG model retrieves external information to generate answers. If the logs show that it directly uses NYT articles without proper citation, each output would be an infringement, and OpenAI cannot rely on “fair use” as a defense.

4. OpenAI’s Knowledge of the Infringement: If the logs prove that OpenAI was aware of the infringement risk but still lied to the court, it will appear dishonest, casting doubt on all its arguments.

The Lessons from Google’s Experience with Evidence Concealment

Previously, Epic sued Apple and Google for monopolistic practices (due to games being sold outside app stores):

  • Google Lost: The court found that Google automatically deleted chat records, determining that this was an attempt to destroy evidence and ruling in favor of Epic’s claim that the deleted content was detrimental to Google. As a result, Epic was awarded a monopoly lawsuit victory.
  • Apple Won: Apple was not accused of destroying evidence, and despite its more closed system, it was not found to be monopolistic.

This shows that tactics used to hide evidence in lawsuits can often determine the outcome more than the facts of the case itself.

What Would Happen if Evidence Were Concealed in China?

Chinese law is also strict:

  • Presumption of Evidence: If one party holds evidence but refuses to provide it, and the other party claims it is detrimental, the court will assume the latter’s version to be true.
  • Penalties: Fines (up to 100,000 yuan for individuals and 500,000 yuan for organizations) and detention (up to 15 days) can be imposed for fabricating or destroying evidence. In severe cases, it may constitute the crime of obstructing justice or assisting in the destruction/fabrication of evidence, leading to imprisonment.
  • Perjury by Witnesses: Witnesses who give false testimony can also face fines and detention.

Honesty in providing evidence is essential in litigation; attempting to outsmart the system can result in significant consequences.

The Ultimate Impact on OpenAI

OpenAI’s defense based on fair use could collapse if it cannot provide the complete logs. Whether the court forces OpenAI to submit the logs or assumes that the concealed content favors The New York Times, OpenAI’s argument will be weak. The outcome of this case will not only affect OpenAI but also serve as a warning to the entire AI industry: using copyrighted material in AI training cannot rely on hiding evidence; legal solutions (such as purchasing rights or using open-source content) must be found. Otherwise, even the most advanced technology can fail due to legal details.