Startups & Funding

Anthropic reveals how Claude's AI watermarks will work

Claude’s new AI watermarks can’t be seen but can always be detected... unless rewritten

Deep Dive

Anthropic has published a blog post detailing how its upcoming AI watermarking system for Claude will function, aiming to comply with the EU AI Act’s Transparency Code. The system, based on Google DeepMind’s 2024 SynthID-Text approach, embeds undetectable patterns in Claude’s responses during low-stakes word choices. These watermarks are detectable only with a specific key and do not compromise output quality—watermarked and unwatermarked text appear identical to readers.

The company addresses concerns about evasion, noting that light editing won’t remove the watermark, while a full rewrite (where every word is replaced) will. For code, watermarks are minimal since the model prioritizes functionality over arbitrary choices, though comments may still carry detectable patterns. Anthropic also plans to release a watermark detection API and clarifies that this method differs from traditional AI detection tools, which rely on identifying stylistic 'tells' in writing.

Key Points
  • Anthropic’s Claude uses Google DeepMind’s SynthID-Text for watermarking, detectable only with a specific API key
  • Light edits won’t remove watermarks; full rewrites will. Code watermarks are minimal but possible in comments
  • Watermarking complies with EU AI Act’s Transparency Code and won’t degrade output quality

Why It Matters

EU-compliant AI watermarking could become standard, forcing transparency in content authenticity and misuse prevention.

📬 Get the top 10 AI stories daily