Catching the Careful
Anthropic will start marking what Claude produces. Text gets an imperceptible watermark woven into the words themselves, and supported files (.svg, .png, .jpg) get signed C2PA provenance metadata. Models launched in the EU on or after 2 August 2026 carry it from launch, older models are being retrofitted during the grace period the law allows, and the marking is applied at the model level worldwide, across the API, Claude Code, Cowork and the cloud partners. None of it is optional and none of it is visible.
The driver is Article 50 of the EU AI Act, in force since 2 August, which requires machine-readable marking of synthetic content that is "effective, interoperable, robust and reliable as far as technically feasible". Anthropic signed the accompanying code of practice, and the code is voluntary while Article 50 is not. Marking obligations are arriving through other doors too. California bundled watermarking guidance into its procurement rules back in March.
The C2PA half is close to decorative. The manifest comes off with a re-save, a format conversion, a screenshot, or an upload to a platform that rewrites metadata on the way in, and The Verge notes it is sometimes stripped by accident. Marks written into the pixels themselves hold up far better. OpenAI's checker was solid enough that people used it two days ago to identify an anonymous model on Arena. Metadata riding alongside a file has never managed that.
The text watermark is the consequential half, and it is the one with nothing published behind it. Anthropic doesn't name the scheme, hasn't released technical documentation, and hasn't shipped the detection tool it says it is working on. The properties it does claim are that the mark rides inside the text, so copy-paste carries it, while heavy editing, paraphrase or translation can leave nothing detectable, as can a passage that is simply short.
Those two properties invert who gets caught. C2PA fails indiscriminately, coming off for anyone who re-saves the file. The text mark is selective, and selective in the wrong direction: someone pushing Claude's output through a second model to pass it off as their own clears it without trying, while someone who pasted Claude's tidy-up of a paragraph they wrote themselves keeps it. The careful user stays marked.
Anthropic is straight about the limits. A detection means content may have been processed by Claude, not that Claude wrote it, and no detection proves a human did. That is the sentence that matters, and it sits in a help-centre article. Detection tools are promised rather than delivered, so nothing has gone wrong yet. Once they arrive, the result will be read out in misconduct hearings and HR meetings by people who never opened the caveat, treated as a test result rather than a hint, the same way plagiarism-checker percentages got treated. Marking the output is the right call. Shipping the marks a long way ahead of any means of interpreting them is not.
Sources:
-
How Claude marks AI-generated content — Claude Help Center
-
Claude will apply invisible watermarks to AI text and images — The Verge
-
Anthropic watermarks all Claude outputs globally — The Decoder
-
Code of Practice on Transparency of AI-generated Content — European Commission
Filed under AI & machine learning
This post is timestamped using Blockchain technology. Verify