Follow Cyber Kendra on Google News! | WhatsApp | Telegram

Add as a preferred source on Google

Anthropic to Watermark All Claude-Generated Text

Anthropic will embed invisible watermarks in Claude-generated text and signed C2PA metadata in generated files, under the EU AI Act transparency code.

anthropic watermark
Anthropic has detailed plans to mark content generated by its Claude models, embedding invisible watermarks directly into generated text and attaching cryptographically signed provenance metadata to generated files.

The company set out the approach in a support document confirming it has signed the EU AI Act's Article 50(2) Code of Practice on Transparency of AI-Generated Content, as a provider of both generative AI models and generative AI systems. 

The document is framed throughout as a description of how Anthropic is "planning to put those commitments into practice", and the company says it will publish more detailed technical guidance as it becomes available.

Claude models launched in the EU on or after 2 August 2026 will support machine-readable marking at launch, according to the company. Anthropic says it is also working to add marking support to models released before that date, which fall under a transition period allowed by the law.

Two marking mechanisms

Anthropic is deploying two techniques that address different failure modes.

The first is an embedded text watermark. According to the company, when a supported model generates text, it weaves an imperceptible mark into the text itself, without changing meaning, quality or readability. Because the mark is carried by the text rather than attached to it, Anthropic says it travels through copy-and-paste and may survive some editing.

The second is signed provenance metadata on generated files. For supported formats including SVG, PNG and JPEG, Claude attaches metadata conforming to the Coalition for Content Provenance and Authenticity (C2PA) open standard. Anthropic says a signed label indicates a file was processed by Claude and allows detection of subsequent tampering.

The distinction matters in practice. Metadata is information-rich but fragile — it does not survive screenshots, format conversion or most CDN image pipelines. Watermarks carry far less information but persist through transformations that destroy metadata. Deploying both is what makes the pair complementary rather than redundant: each covers the other's principal failure mode.

Applies worldwide, across every surface

Although the commitments originate in EU law, Anthropic states that marking will apply to output from supported models wherever Claude is offered, worldwide.

Coverage spans the Claude Platform API, the Claude consumer apps, Claude Code, Claude Cowork and Claude Tag. Embedded watermarks will also apply when supported models are accessed through AWS, Google Cloud and Microsoft Foundry, though Anthropic notes that signed provenance metadata may not be supported on every cloud platform, depending on the features each offers.

Anthropic says watermarking is applied at the model level, meaning it is present regardless of which product or surface the text comes from.

Detection remains unavailable

The company says it will support users and third parties in detecting Claude's marks, as the Code requires, with details to follow in forthcoming technical documentation.

Until those ships, the text watermark is unverifiable by anyone outside Anthropic. Keyed text watermarking schemes bias token selection during generation according to a secret; confirming the signal requires that secret. Google's SynthID text watermarking runs on Gemini output, but its detector remains restricted to Google and selected enterprise partners, and OpenAI has not deployed a detectable text watermark despite publishing research on the technique.

The practical consequence is a gap between marking and verification. Once marking is live, output will carry a signal that no publisher, platform or academic institution can currently read.

The C2PA file metadata is a different matter — it uses an open standard with public tooling, so anyone can verify it today using the C2PA verification tool, c2patool, or any C2PA-compatible reader.

Anthropic's own caveats

The support document is unusually direct about what marks do not establish.

A detected mark indicates content may have been processed by Claude, but it is not conclusive. Anthropic notes that people routinely use Claude to proofread, translate, summarise or convert files, so output can carry a mark even when the underlying ideas or text originated elsewhere. Marked content may also be edited, excerpted or combined with other material afterwards.

The reverse holds more strongly. Anthropic lists several conditions under which genuinely AI-generated content will carry no detectable mark: generation by a model predating marking support, heavy editing or paraphrasing, translation, passages too short to carry a reliable signal, metadata stripped through format conversion or re-saving, and platforms or file types where a marking type is not supported.

For anyone building detection into an editorial or compliance workflow, that asymmetry is the operative fact. A mark found is evidence. A mark absent is nothing at all.

What it means for publishers

Organisations deploying Claude in their own products carry independent obligations. Anthropic says those building with Claude should assess what Article 50 requires of their own products and services, and that it will publish technical guidance to support those obligations.

For publishers, the near-term implication is narrower than it first appears. Text watermarks cannot be checked by anyone yet. File credentials can be, and any newsroom accepting image submissions can start verifying them today — a valid credential identifies the signing tool, the certificate holder, and the exact time of signing, and proves whether the file has changed since.

Post a Comment