Anthropic Adds Invisible Watermarks to Claude's Text and Images to Meet Provenance Rules
Copy a paragraph out of Claude, paste it into an email, a term paper, or a marketing brief, and for the past four years the connection between those words and the machine that produced them simply evaporated. Anthropic wants to change that. In an announcement made around August 10, the company said it has begun weaving machine-readable watermarks directly into text generated by supported Claude models, plus digitally signed provenance metadata for images and other files. The mark, Anthropic says, is imperceptible to a reader, does not degrade the response, and can travel with the text even after it is copied, pasted, and lightly edited.
It is one of the first large-scale attempts to make AI-generated prose carry a provenance trail the way a photograph carries EXIF data, and it lands squarely at the intersection of regulation, academic integrity, and a detection industry that may soon need to reinvent itself.
How the watermark works, and where it stops
Anthropic is using two distinct mechanisms. Text receives an embedded watermark applied at the model level, which means it appears regardless of which product generated it. Supported files in PNG, JPG, and SVG formats instead carry digitally signed provenance metadata built on the C2PA standard, the open specification maintained by the Coalition for Content Provenance and Authenticity that records how a file was created and whether it has been altered.
"When a supported Claude model generates text, it weaves an imperceptible watermark directly into the text itself," Anthropic wrote in its support documentation. "You won't see it, and it doesn't change the meaning, quality, or readability of Claude's response." Because the signal lives inside the words rather than in file metadata, the company adds that it "will travel with the text when it's copied and pasted elsewhere, and may persist through some editing."
The coverage is deliberately broad. Anthropic says the marks apply across its API, the Claude apps, Claude Code, Claude Cowork, and Claude Tag, and extend to supported models accessed through AWS, Google Cloud, and Microsoft Foundry. This is not a label bolted onto one chat window; it is baked into model output at the source.
The limits, though, are as important as the capability, and Anthropic is unusually candid about them. A detected watermark indicates only that Claude may have processed the content, not that Claude authored it. A human could write an essay, run it through Claude for proofreading or translation, and end up with marked text despite owning every idea in it. The reverse also holds: the absence of a mark does not clear a piece of writing, because heavy editing, aggressive paraphrasing, translation, mixing Claude output with other text, or passages that are simply too short can all weaken or erase the signal. Older Claude models may not support marking at all, and image provenance can vanish through a screenshot or a format conversion.
Independent researchers reinforce that the signal is fragile. Jonas Geiping, who leads a machine learning safety group at the ELLIS Institute Tübingen and studies watermarking, said that paraphrasing can remove a watermark but that not every paraphrase will; stripping one out of a long document, by his account, takes more than a light pass because enough of the original phrasing has to change. Alex Cui, CTO and co-founder of the detection firm GPTZero, published a technical explainer arguing that text watermarks can be defeated, noting that free tools have already bypassed Google DeepMind's SynthID. Notably, Anthropic has not yet published the details of its own detection mechanism, so those tests describe the general class of technique rather than Claude's specific implementation.
The regulatory engine behind it
The policy stems from Anthropic signing the EU AI Act's Article 50(2) Code of Practice on Transparency of AI-Generated Content, which obliges providers to make synthetic output machine-readable and detectable. Claude models launched in the EU on or after August 2, 2026 support marking at launch; earlier models fall under a transition period Anthropic says it is still working through, without committing to a date.
Anthropic is far from alone. The European Commission counted roughly 190 signatories to the code by the end of July, and Google, Meta, Microsoft, Mistral, and OpenAI have all joined the provider section covering marking and detection. What makes Anthropic's rollout striking is its reach: the company is applying the marks worldwide, not just to European users, effectively letting an EU compliance obligation set a global default.
The move also underscores how far the industry's posture has shifted. As Search Engine Journal noted, OpenAI scrapped its own ChatGPT watermarking plans in 2024 after an internal survey found that nearly 30 percent of users said they would use the product less if watermarking were added. Regulation appears to have overridden that reluctance.
Why it matters for schools and the detection business
The implications land hardest in two places. The first is academic and editorial integrity. For institutions whose standard is that submitted work be entirely human-written, a Claude mark could function as a red flag even though Anthropic explicitly warns it does not prove authorship. A translator working from someone else's article, or a student who ran a self-written draft through Claude for grammar, would both trip the signal. Any honor code that treats a positive hit as proof of cheating inherits that ambiguity, and the EU's own Article 50 carves out exemptions for standard editing and for text under genuine human editorial control, meaning a mark can appear on copy no one is obligated to disclose.
The second is the AI-detection industry itself. Companies like GPTZero built businesses on probabilistically guessing whether text came from a machine. A deliberate, verifiable signal placed by the model maker is a different animal, and if OpenAI, Google, and others follow with interoperable marks, detection could shift from statistical inference to checking a registry of provenance signals, forcing today's detectors to become provenance platforms instead.
What to watch next
The decisive unknown is detection. Anthropic says it will give users and third parties a way to verify its marks and will publish technical documentation, but it has not said what that access will look like, and public detection cuts both ways: the same tools that let a teacher check an essay give a determined evader something to test their removal tricks against. Watch for whether Anthropic keeps verification inside its own products, as Google has largely done with SynthID, or opens it up. Watch, too, for how quickly the other 190 signatories ship compatible systems, and whether any court or university treats a watermark hit as evidence. Until the detection layer is real and interoperable, Anthropic's invisible ink is a promising signal in search of a reader.
“When a supported Claude model generates text, it weaves an imperceptible watermark directly into the text itself. You won't see it, and it doesn't change the meaning, quality, or readability of Claude's response.”— Anthropic, Support documentation, 'How Claude marks AI-generated content'