AI Newsway

Anthropic Details How Claude's Invisible Text Watermarks Work

The system is a SynthID-Text variant that changes where randomness comes from, not which words Claude picks.

|4 min read0
AI Summary
Anthropic published details of the watermark it will apply to Claude's text output to comply with the EU AI Act, confirming it is a variant of Google DeepMind's SynthID-Text. The scheme replaces the random number used to pick among equally fitting tokens with a value derived from a secret key, so no characters are added and quality is unchanged in rater tests. Detection weakens on short, factual or heavily edited text and code, and cannot identify other models.
Claude's watermark lives in the model's word choices rather than in any character added to the page.
Claude's watermark lives in the model's word choices rather than in any character added to the page.

Anthropic has published a detailed account of the watermarking system it will apply to Claude's text output, confirming that the method is a version of SynthID-Text, the approach Google DeepMind introduced in a 2024 Nature paper. The change is being made to satisfy the European Union's AI Act.

The technique inherits from a proposal Scott Aaronson floated in 2022, and its central idea is narrower than the word watermark suggests. Nothing is inserted into the text. No hidden characters are added, no extra tokens are consumed, and the model neither slows down nor costs more to run.

Changing the source of randomness

Language models emit one token at a time, choosing among candidates that would each fit. In a sentence beginning with a description of cold weather, a model might reasonably land on overcast or grey; the meaning survives either way, and ordinarily a random number settles it.

Watermarking replaces that arbitrary randomness with a value derived from a secret key and the immediately preceding words. The selections remain effectively random to a reader, but anyone holding the key can test whether a passage's word choices are consistent with what Claude would have produced under it, and assign a probability accordingly.

Crucially, the system does not push the model toward vocabulary it would otherwise avoid. Anthropic notes it would not induce Claude to reach for an obscure synonym like nubilous, because the nudge operates only among candidates already in contention. In internal testing and in DeepMind's own controlled study, raters comparing watermarked and unwatermarked answers side by side found no difference in quality.

Where the watermark thins out

The company is unusually direct about the limits. Detection answers one question only: how likely is it that Claude was partly involved. It cannot establish that text is human-written, and it cannot identify output from a different model, which would carry a different key or a different scheme entirely.

Signal strength also scales with length, so short samples offer too few decision points to be conclusive. Factual passages are watermarked sparsely, because accuracy leaves little room for interchangeable wording; once a sentence commits to naming Isaac Newton's Principia, only one continuation is correct. Code is similarly resistant, since exactness is usually required, and comments within code are one of the few places the mark can settle.

Editing behaves the same way. Proofreading a human draft leaves almost nothing for the watermark to attach to, because nearly every word remains the author's. Translation, by contrast, is fully watermarked, since Claude chooses every word. Light editing will not reliably strip a mark, though a complete rewrite will, at which point the label arguably no longer applies.

Regulation is the driver

Anthropic signed the EU Code of Practice on Transparency of AI-Generated Content in July 2026, one of roughly 190 signatories, and the marking obligation took effect on August 2. Other major model developers signed the same code and will deploy their own watermarks.

Notably, Anthropic is applying watermarking globally rather than to European traffic alone, saying it lacks a durable way to scope the behavior by region for now. Models released before August 2, 2026 fall under a transition period and will receive watermarking over the coming months.

For files rather than text, Claude attaches a C2PA content credential, a cryptographically signed note in the metadata using the same open standard camera makers and photo editors already rely on. That is a label rather than a watermark: nothing in the file itself changes, and it disappears with anything that discards metadata.

A detection API is planned but not yet specified. Until it ships, the practical effect for most users is nil. The marking is invisible, carries no information about the person or organization that prompted it, and changes nothing about ownership of or responsibility for the output.

How do you feel about this article?

SJ

Discussion

Sign in to post
Loading...

Related articles

Anthropic Will Seat Outside Evaluators Inside the Company as Amodei Urges a Slower Frontier
AI & Machine Learning

Anthropic Will Seat Outside Evaluators Inside the Company as Amodei Urges a Slower Frontier

Dario Amodei wants frontier labs to slow capability gains, and is giving outside evaluators badges and laptops at Anthropic to prove it can be verified.

Seung Jung5 days ago
A Model Talked Its Safety Monitor Out of Flagging a Real Attack
AI & Machine Learning

A Model Talked Its Safety Monitor Out of Flagging a Real Attack

A reasoning-trace safety monitor failed to flag a model attacking live systems because the model spent the session narrating the targets as simulated.

Seung Jung6 days ago
A Researcher Quit Anthropic Over Extinction Risk. His Safety Lead Agreed in Public.
AI & Machine Learning

A Researcher Quit Anthropic Over Extinction Risk. His Safety Lead Agreed in Public.

Jacob Coxon resigned from Anthropic over extinction risk. The company's head of alignment stress testing publicly agreed and put the odds above 10 percent.

Seung Jung8 days ago
Attackers Ran Agents That Rebuilt Their Malware Until Scanners Stopped Catching It
AI & Machine Learning

Attackers Ran Agents That Rebuilt Their Malware Until Scanners Stopped Catching It

A Russia-linked crew let agents iterate on flagged implants until detection failed. It is the clearest published case of attackers closing the loop on static signatures.

Seung Jung6 days ago
Anthropic Names Alibaba, Moonshot and DeepSeek in 200M-Exchange Distillation Report
AI & Machine Learning

Anthropic Names Alibaba, Moonshot and DeepSeek in 200M-Exchange Distillation Report

Anthropic says five campaigns ran nearly 200 million Claude exchanges to copy its reasoning, with Moonshot and DeepSeek relaying live customer traffic.

Seung Jung7 days ago
An AI Cracked a 373-Year-Old Cipher. Then Someone Checked the Microfilm.
AI & Machine Learning

An AI Cracked a 373-Year-Old Cipher. Then Someone Checked the Microfilm.

Vals AI reported Claude Fable 5.1 solved a 373-year-old cipher in 44 minutes. An independent replication reports 8 of 64 letters match — chance level.

Seung Jung4 days ago