Summary of Key Points
Anthropic (the parent company of Claude) has added invisible AI watermarks to all the text, images, and files generated by Claude in response to EU regulations and to gain a foothold in the industry. However, these watermarks were cracked within just three days by an open-source tool on GitHub called watermark-remover. This tool can remove statistical patterns from the text, strip metadata from files, and even regenerate images to eliminate pixel-based watermarks. Behind this “watermark battle” lie the conflicting needs of users: they want to be able to identify AI-generated content to prevent fraud, but at the same time, they fear that their works created with AI assistance will be labeled as “fully AI-generated,” thereby denying the value of their labor.
I. Claude’s Watermarks: The Invisible “Tattoos” Hidden in the Content
Claude’s watermarks are not visible logos or texts; rather, they are statistical patterns embedded in the generation process.
When AI generates text, it’s like a word-game where each word is chosen based on the highest probability from a list of candidates (for example, “growth rate is very XX,” with possible options being “fast,” “rapid,” or “astonishing”). After adding the watermarks, Claude subtly adjusts the probabilities of certain words (e.g., increasing the likelihood of “fast” and decreasing that of “rapid”). Individually, these changes are not noticeable, but when accumulated in a long piece of text, they create a “code” that only machines can recognize.
For images and files, the approach is more direct: images are tagged with digital signatures, and PDF/Word documents contain hidden source identifiers. Anthropic claimed that these watermarks could not be copied and pasted or removed with simple editing, but this claim was quickly disproven.
II. The GitHub Crackdown Tool: Cleaning AI Content Completely
The power of watermark-remover lies in its comprehensive coverage:
1. Text Watermarks: It either removes hidden characters (such as Unicode wide spaces) or uses another AI model to rewrite the text, disrupting the statistical patterns (at the cost of potentially losing the original tone and style).
2. Image Watermarks: It uses the CtrlRegen tool to “reconstruct” the image, retaining the content while generating new pixels, effectively erasing the watermark.
3. File Watermarks: It removes metadata like EXIF/XMP from PDF/Word files (equivalent to tearing off the file’s “identity card”).
4. Self-Verification: The tool integrates with a detection framework from Tsinghua University to confirm that the watermarks have indeed been removed.
Ironically, Claude itself refuses to run this tool—even when users request it, claiming, “I wrote the content; I just want the watermark removed,” implying, “The marks I added, I can’t take away.”
III. Why the Battle? Regulatory Pressure, Corporate Ambition, and User Resistance
This conflict is not incidental:
1. Regulatory Push: Over 190 AI companies have signed the EU’s “AI Generated Content Transparency Guidelines,” requiring AI-generated content to include machine-readable markers. Anthropic was one of the first to comply, aiming to gain control over the discussion on “AI safety.”
2. Corporate Ambition: The company that develops advanced watermark technology will have a competitive advantage in future industry standards (Anthropic used Google DeepMind’s SynthID solution to showcase its technical capabilities).
3. User Resistance: Watermarks are seen as a label of “original sin” for AI-generated content. Users argue that if they spend hours researching and formatting their work, and then use Claude to refine it just two sentences, their efforts should not be overshadowed by the same watermark as completely AI-generated content. They fear that their labor is dismissed as merely AI-generated.
IV. The Core of User Disapproval: Watermarks as a One-Size-Fits-All Label
Users are not opposed to marking AI content in general, but they dislike the brutal nature of these watermarks:
- Blurred Boundaries: There is no clear distinction between AI-assisted creation (e.g., polishing or researching) and fully AI-generated content; watermarks treat both as if they were the same.
- Denial of Value: Users worry that watermarked content may be restricted by platforms or questioned by readers, dismissing the value of their efforts.
- Lack of Choice: Claude automatically adds watermarks to all users, regardless of their preferences, which they feel infringes on their control over their own work.
V. Is This Battle Without End? Need for Smarter Regulations
The battle between watermarking and removal tools will continue: as Anthropic upgrades its watermarking techniques, the open-source community will develop more powerful removal tools. The real issue is not technology itself, but how to balance the need to identify AI-generated content with the protection of creators:
- Can we implement graduated labeling (e.g., distinguishing between “AI-assisted” and “fully AI-generated” content)?
- Can users make choices (e.g., paying to have watermarks removed)?
- Can regulations be more flexible (e.g., requiring watermarks only for high-risk AI content like news or ads, while allowing them to be optional for other types of content)?
If we remain in a cycle of “add watermark, remove it, add it again,” the ultimate victims will be ordinary users—either falling victim to AI scams or having their work wrongly labeled. Perhaps watermarks will become obsolete when AI-generated content is indistinguishable from human-created content, but that day is still far off.
This battle is essentially a question of identity in the AI era: we need to prevent AI from impersonating humans while ensuring that the benefits of AI are recognized. Current watermarking technologies have not yet found this balance.