虎嗅

Claude’s text watermark successfully outsmarted human strategies.

原文:Claude 的文字水印,成功把人类套路了

Summary of Key Points

Anthropic added an invisible watermark to the text generated by Claude, utilizing Google SynthID. The principle behind this watermark is to subtly adjust the probabilities of word selection during generation, creating statistical patterns that only detection tools can recognize. However, anti-watermarking tools (such as those used for AI-based content moderation) have emerged quickly and can easily break these patterns. This technological battle not only exposes the vulnerability of watermarks but also raises a deeper issue: when both text creation and modification are aimed at avoiding detection or leaving traces, the pursuit of word precision and style in writing is being overshadowed by technical objectives.

1. Claude’s Invisible Watermark Isn’t Just a “Stamp”

Do you think a watermark is just a hidden symbol added after the article is written? Not quite. When Claude generates text, it calculates the probabilities of each potential word (or token) before selecting one. For example, for the phrase “This movie is really ____,” the possible options might be “good-looking” (40%), “impressive” (20%), and “exciting” (10%). SynthID then modifies these probabilities according to a set of secret rules—perhaps increasing the likelihood of “good-looking” to 45% and decreasing “impressive” to 18%—to influence the model’s choice.

However, it doesn’t force the model to use a specific word, nor does it have a fixed list of “watermark words.” The same word might be favored in one context but not another. Individually, there’s no obvious anomaly; only after hundreds of words are generated can detection tools identify that the word selection follows a particular statistical pattern. It’s like a student consistently choosing option C on exams; it’s not coincidental but rather a result of some form of “guidance.”

2. Why Are Anti-Watermarking Tools So Easy to Use?

The reason anti-watermarking tools are effective is that the watermark is closely integrated with the text itself. It’s not hidden in file metadata (like image shooting information) or represented by invisible characters; instead, it’s directly embedded in the word selection process. To remove the watermark, you simply need to reword the text without changing its meaning—use synonyms (“good-looking” becomes “exciting”), change the verb tense (active to passive), or use another AI tool for moderation (like Declaude, which is designed to remove artificial elements). It’s similar to avoiding a teacher’s preference for certain words in an essay; you just need to replace those words with others that convey the same meaning, and the teacher’s influence disappears. The cost of countermeasures is extremely low, which is why Claude’s watermark proved problematic almost immediately.

3. Asymmetric Battle: Cheaters Can Escape, While Honest Users Get Marked

This situation is unfair to two types of people:

  • Cheaters: They will actively use non-watermarked models or content moderation tools to repeatedly modify the text until it can no longer be detected.
  • Ordinary Users: They may not be aware of the watermark and use Claude by default, only to find their content marked as AI-generated, potentially leading to misunderstandings.

It’s like in exams where cheaters prepare in advance to avoid the teacher’s guidance, while well-prepared students end up being flagged for choosing the “correct” answer (in this case, option C) according to the teacher’s hints. The easiest targets for anti-watermarking tools are those who don’t try to hide anything.

4. The Biggest Threat Isn’t Technological Failure, but the Hijacking of Language Expression

Anthropic claims that the watermark doesn’t affect the meaning or quality of the text. However, the real issue is that when models choose words, they now have an additional goal: whether the choice will be detected or not. This shifts their focus from “how to write more accurately and stylishly” to “how to avoid detection.” Users also modify their wording to evade detection, ignoring subtle differences between words.

For instance, phrases like “He finally agreed” and “He finally relented” have similar meanings, but “relented” implies a sense of resistance and negotiation. For those who truly care about writing, these words cannot be exchanged at will. Yet now, both models and users use the idea of “similar meaning” as a justification for making changes. Language is no longer used for precise expression; instead, it has become a pawn in technical battles.

In the end, the debate isn’t about whether watermarks are visible or not, but whether the pursuit of style, rhythm, and precision in writing can still be valued. When everyone is modifying their words to fit technical requirements, the original intention behind the text is being marginalized.

Conclusion

Although Claude’s watermark technology seemed advanced, it was quickly countered. What’s more concerning is that it has turned language into a battlefield for technological conflicts—language no longer serves its purpose of conveying meaning but must now consider whether it will leave traces or avoid detection. This highlights a critical point: as technology evolves, we must not forget the intrinsic value of language itself. After all, the choice of words is never about simply conveying similar meanings; they are an integral part of the meaning itself.