虎嗅

Claude Announces the Use of Invisible Watermarks for Text: No AI-generated Content Can Hide Its Origins Anymore

原文:Claude 宣布给文字打上隐形水印,你用AI 写的每个字都藏不住了

Summary of Key Points

After the new EU Artificial Intelligence Act came into effect, requiring generative AI to add detectable markers to its content, Claude, a product under Anthropic, was among the first to introduce two types of "invisible" markers: embedded text watermarks and file metadata signatures. These measures aim to address the challenge of distinguishing between authentic and AI-generated content (for example, AI-written articles being passed off as original or original works being mistakenly identified as AI creations). However, watermarks are not omnipotent; they have limitations and have led to new issues such as the emergence of a business service that reduces the perceived "AI content ratio" in documents. Ultimately, the value of watermarks lies not in convicting AI-generated content but in increasing the cost of fraud and making it clearer who is responsible for the authenticity.

Why Add Watermarks to AI Text?

The two most frustrating issues for ordinary people are:

1. AI-generated content being presented as if it were manually written (for instance, an AI-written article claiming the author worked all night without any evidence to prove otherwise).

2. One's own hard work being questioned as possibly AI-generated, with no way to prove its authenticity.

The EU regulation sets a clear requirement: generative AI must add machine-readable markers to its content, with fines of up to 15 million euros or 3% of global turnover for violations. This is in response to real-world problems such as:

  • Second-hand platforms selling AI translations of Haruki Murakami's works without the author's latest publications, leaving buyers unaware that they are AI-generated.
  • A book titled "Documentary on Fan Culture" being questioned as potentially AI-written, with the author's field research failing to convince everyone of its authenticity.
  • Universities mistakenly identifying a poem as AI-generated using outdated detection tools.

Watermarks serve to provide a way of verifying the origin of AI content, helping users determine whether the text came from an AI model and reducing disputes based on subjective assumptions.

How Does Claude Apply Watermarks?

Claude uses two methods that do not affect the readability of the content:

1. Embedded Text Watermarks: These are like invisible tattoos embedded in the text. The signals are hidden within the language and sentence structure; humans cannot see or detect them, but machines can. These watermarks remain even when the text is copied, pasted, or slightly edited, whether generated through the Claude app or its API.

2. File Metadata Signatures: For image files (such as SVG, PNG, or JPGs), a unique identifier is added to verify that they have been processed by Claude. For example, if you use Claude to create a chart, the file will contain this signature upon download.

Users or third-party tools can detect these markers, but the technical details are still to be revealed in official Anthropic documentation.

Watermarks Are Not a Panacea

Watermarks have several limitations:

  • They cannot definitively prove that all content is AI-generated. For example, if you use Claude to revise your own article and the output includes a watermark, the core content remains your original work.
  • Significant modifications to the text may make the watermark invisible. If the text is heavily edited or translated back, or if paragraphs are too short, the watermark signals could be lost.
  • Old content or files with changed formats may lose their watermarks. For instance, Claude models prior to August 2nd did not have watermarks, and screenshotting or converting file formats (e.g., to PDF) could remove them.
  • The absence of a watermark does not necessarily mean the content is not AI-generated. Content created by other models (like older versions of ChatGPT) or thoroughly cleaned may also evade detection.

Consequences of Watermarks

The implementation of these regulations has led to direct changes in real-life situations:

  • Students' Struggles: Universities have linked the proportion of AI-generated content to grading requirements, with varying detection results between tools. To pass, students are altering their papers to make them more conversational in style and even purchasing services to reduce the perceived AI content ratio (with sales exceeding 4,000 such services).
  • Employees' Dilemmas: Companies train employees to use AI for efficiency but require that the proportion of AI-generated content does not exceed a certain limit. Employees often have to rewrite AI-generated parts manually, which is more time-consuming.
  • The Emergence of a New Business Niche: Services that reduce the perceived AI content ratio are becoming popular on e-commerce and social media platforms, costing anywhere from tens to hundreds of euros per document.

The True Value of Watermarks

The real value of watermarks is not in convicting AI-generated content but in increasing the cost of fraud:

  • They make it more difficult to create large quantities of fake AI content, as each piece will have a recognizable marker, making it easier to catch.
  • In the future, watermarks might become less of a concern as they become commonplace. People may start to find content without them more suspicious.
  • Watermarks do not answer the fundamental question of who is responsible for the content. Whether an article uses AI or not, readers generally do not mind as long as the facts are verified and someone is willing to take responsibility.

In summary, while watermarks are not a perfect solution, they mark the beginning of better management of AI-generated content, making it easier to trace its origin and clarify responsibilities.

(End of translation)