Anthropic Explains Claude Watermark Mechanics
Anthropic details how Claude’s watermarking works, addressing evade, edit, and code concerns.
Anthropic published a blog post Friday answering basic questions about its plan to watermark text generated by its Claude chatbot, a move that has sparked debate among users. The company is adopting the approach to comply with the EU AI Act’s Transparency Code, which requires AI companies to make AI-generated content identifiable. In the post, Anthropic explains how the watermarking works, whether it can be removed by editing, and how it handles code.
The timing is notable. Earlier this week, Anthropic revealed its watermarking plans, and the announcement immediately drew a strong reaction from users. On Reddit, one poster characterized this as a conspiracy against innocent Claude users, while another claimed, “The only reason you wouldn’t want this is to lie to people.” And Business Insider reports that “dozens” of users on X have claimed to cancel their Claude subscriptions as a result.
The new post is Anthropic’s attempt to provide clarity. It starts with a general overview of the concept, explaining that when making “low-stakes choices” — like choosing between the words “overcast” and “grey” to describe the weather — Claude can create a pattern in its responses that is “undetectable to the reader, but is detectable to anyone who has a key that encodes it.”
“Watermarking does not impact the quality of Claude’s output. To a reader, a watermarked response is indistinguishable from an unwatermarked one.”
— Anthropic, in its blog post
How the Watermark Works
Anthropic said it will use the SynthID-Text approach that the Google DeepMind team outlined in 2024. The company also plans to release a watermark detection API, though it didn’t provide specifics on when that API will be available or how developers will access it.
The company emphasized that watermarking is distinct from AI detection approaches offered by companies like Pangram, which look for stylistic “tells” in the writing (like the construction “his isn’t [X], it’s [Y]”) to reveal AI usage. “Picking up on these patterns is fundamentally different from checking for a watermark,” Anthropic said.
Can the Watermark Be Hidden by Editing?
One of the most pressing questions is whether users can simply rewrite or edit the text to remove the watermark. Anthropic acknowledged it’s possible, but “light editing probably won’t remove the watermark completely,” while “a complete rewrite where every word is replaced will.”
“In the latter case, of course, it’s arguable whether the text can any longer be described as AI-generated,” the company said.
What About Text Claude Proofreads or Edits?
For users who use Claude to proofread or edit their own writing, the watermark situation is more subtle. Anthropic said it will depend on “the length of the text and how heavily Claude has edited it.” If it’s only been lightly edited, “nearly all the words” will have been written by the human author and “there’s very little (if anything) for the watermark to attach to.”
This suggests that the watermark is more likely to be present in text that Claude generates from scratch, rather than text that a human has substantially authored and Claude merely tidied up.
The Impact on Code
Code will have less of a watermark than other text, because the model needs to create working code and won’t have the freedom to choose between a variety of equally valid options. Anthropic said that “in areas where there is an arbitrary choice between particular words or terms within the code, the watermark can be used, such as comments within code.”
“But by definition, it will have a negligible effect on the actual code produced,” the company said.
Beyond Claude
Anthropic also noted that Claude won’t be the only AI chatbot to generate watermarked text. The company said that “other major model developers have signed the same Code of Practice and will be implementing their own watermarks.”
This suggests the development is part of a broader industry response to regulatory requirements, not an isolated move by Anthropic.
Why It Matters
These details could shape how users and businesses rely on AI-generated text going forward. For businesses, the ability to verify AI-generated content could become a standard practice, especially in regulated industries where transparency is required. For individual users, the watermark may raise concerns about privacy and control over their own content, even if Anthropic says it won’t affect output quality.
The fact that other major developers are expected to follow suit suggests that watermarking could become a common feature across AI chatbots. This could mean that AI-generated text will become easier to identify, but it also raises questions about how effectively the watermark can be preserved in real-world use, especially when text is edited or paraphrased. As the technology evolves, the balance between transparency and usability will likely remain a point of contention.
Sources
- TechCrunch Original source
Continue Reading
Google lets users drop visible AI watermarks
Google will let users remove visible watermarks from AI generations while keeping invisible SynthID and C2PA metadata.
OpenAI's Computer History logs your keystrokes
OpenAI's Computer History feature records clicks and typing, storing them unencrypted and expanding prompt injection risks.
Open AI Weights Draw Defense From Three AI Pioneers
At Ai4, Hinton, Li, and Ng argue for openness in AI despite safety worries, disagreeing on tactics.