Claude's Invisible Ink Marks Involvement, Not Authorship
I reckon the most telling thing about Anthropic’s new watermark is not that it exists, but how modestly the people who built it describe it.
From August, new Claude models will start weaving an invisible mark into their text, a move Anthropic says is to comply with the EU AI Act alongside several other major providers.1 Instead of using an arbitrary random number generator to pick the next word, the system uses the key and a few words that come before to settle what word the model should pick, leaving a pattern you cannot see but a detector with the key can.1 And because marking occurs at the model level, it follows output wherever supported Claude models are offered worldwide, across the API, Claude, Claude Code and the rest, not just inside Europe.2
That sounds grouse until you ask what the pattern actually proves. The watermark can only flag that Claude was likely involved in creating a text. It cannot tell whether Claude wrote the whole thing or just edited it heavily, and it cannot determine whether a text came from a human or a different AI model.3 When Claude proofreads text written by a person, what it gives back has generally only been lightly edited, and because nearly all the words are the person’s, there is very little for the watermark to attach to.4 Anthropic acknowledges the same boundary from the other direction, a watermark does not prove Claude wrote the content, since people often use the AI to edit or translate their own text, and heavy editing or format changes can also strip the markings entirely, however it may persist through some editing.2 The signal is thinnest where you might most want it. Watermarking is sparser on factual passages where there are fewer choices that can be made without decreasing the accuracy of the text.5 These kinds of arbitrary choices do not come up as often with code, so it is less likely to have watermarks at all.6
The generous reading is that this is deliberately light touch. To a reader, a watermarked response is indistinguishable from an unwatermarked one, with the company saying watermarking does not impact the quality of Claude’s output.7 When a supported Claude model generates text, it weaves an imperceptible watermark directly into the text itself, you will not see it, it does not change the meaning, and it travels with the text when copied.8 That is a different kind of evidence from the old detectors. A watermark detector instead looks for a signal deliberately embedded during generation, whereas conventional detectors infer AI authorship from writing patterns, often without access to the generating model.9 The less generous reading, and the one I hear most from practitioners, is about who gets to read that signal. If detection requires access to the base models, it makes Anthropic “the player and the referee” in its own market, the only entity who gets the say on whether a piece of text came out of Claude.10 The practical gap is already visible. Anthropic has not shipped a public Claude watermark detector yet, and the attack path watermark research points to is a meaning-preserving paraphrase with a non-Claude model.11 Until public technical results establish how the system performs across languages, model settings and specialised outputs, the claim that the watermark does not change the meaning, quality or readability of Claude’s text remains a company claim.9 The law Anthropic cites does not help clarify the inference either. Article 50(2) requires providers to mark outputs in a machine-readable format and to ensure that those outputs are detectable as artificially generated or manipulated.12
I do not buy the watermark as an authorship test, and Anthropic does not either. What it gives you is a clue about involvement, not a verdict on who wrote the sentence, who fixed the commas, or who ran it through translation. The gap matters because downstream readers, publishers, universities and regulators are primed to treat any flag as proof, and this system is designed to be invisible to them while legible only to the vendor’s key. The sensible way to treat it is as probabilistic provenance, useful if you already suspect Claude was in the loop and near useless for proving a negative or apportioning how much of the work was the model’s. If you want a mark that distinguishes drafting from light editing or translation, you need metadata that says so, and this one deliberately does not. That is not a failure of engineering, it is the boundary of what low-stakes word choice can carry, and pretending otherwise will cause more harm than the watermark prevents.
Sources
- 1 How Claude’s text watermarking works — anthropic.com
- 2 Anthropic watermarks all Claude outputs globally with marks that “may persist through some editing” — the-decoder.com
- 3 Anthropic announces watermark detection API that will let third parties detect Claude’s AI texts — the-decoder.com
- 4 @joshkel on Anthropic’s ‘watermark’ text adulteration in Claude is a perversion of writing — ycombinator.com
- 5 Anthropic says text watermarking scheme relies on inconsequential words — theregister.com
- 6 @jweber123 on Anthropic’s ‘watermark’ text adulteration in Claude is a perversion of writing — ycombinator.com
- 7 Anthropic shares more details about how Claude’s new watermarks will work | TechCrunch — techcrunch.com
- 8 Anthropic Posts ‘How Claude Marks AI-Generated Content’ Without Explaining How — daringfireball.net
- 9 Anthropic Google Rivalry Moves Into AI Text Watermark Detection — remio.ai
- 10 @pibaker on Anthropic’s ‘watermark’ text adulteration in Claude is a perversion of writing — ycombinator.com
- 11 @arcfour on How Claude marks AI-generated content — ycombinator.com
- 12 Watermarking AI Content Under Article 50(2) AI Act | Stibbe — stibbe.com
How this was made
- 01-research z-ai/glm-5.2 $0.389
- 03-annotate z-ai/glm-5.2 $0.198
- 04-nominate deepseek/deepseek-v4-pro $0.036
- 05-select google/gemini-3.7-flash $0.003
- 06-write meta/muse-spark-1.2 $0.034
- 08-visualise anthropic/claude-sonnet-5 $0.093
total $0.752