Marking is not accountability
For the past three years, every time someone asked how we were going to tell text written by a person apart from text generated by an AI, the answer arrived wrapped in the same promise: we’re working on it, give it time. The solution was called watermarking. An invisible mark, unreadable to the naked eye but legible to a machine, that would travel glued to the text like a fingerprint and would eventually be available to anyone who wanted to check where what they were reading came from.
On August 11th, Anthropic delivered on that promise. Partially, and three years late.
A mark that says less than it promises
What Anthropic announced is, technically, reasonably elegant. Every time Claude generates text, the model chooses among several near-equivalent words to complete a sentence; watermarking biases that choice according to a secret key and the words that precede it, leaving a statistical pattern that only someone holding that key can detect. It’s not hidden Unicode characters or metadata bolted onto the end: the mark is woven into the word choices themselves, applied at the model level — so it travels with the text regardless of whether it comes out of the Claude app, the API, Claude Code, or Claude Cowork — and survives, Anthropic says, some amount of editing afterward. For images, the company leans on C2PA, the standard that cryptographically signs a file with who created it and with which tool.
That’s the part that can be announced without caveats. The caveats came four days later, on August 15th, when Anthropic detailed its own limits: detection fails on short passages, is unreliable when the content is factually constrained — there are only so many ways to write a date or a chemical formula — and becomes useless when the follow-up editing is minimal, because too few watermarked words survive to accumulate a reliable signal. And, above all, it admitted what any student had already suspected: a complete rewrite, word by word, erases the mark without a trace.
What’s left, then, isn’t an answer to “did a person or an AI write this?” It’s what explainx.ai’s technical breakdown summed up with precision: “a volume filter, not a judgement.” The mark doesn’t distinguish between someone who auto-generated a thousand spam articles and someone who asked Claude to polish a paragraph of their own writing. It only says that Claude, at some point, had a hand in it.
Anthropic had already promised this. Twice.
None of this is new, and that’s the part August’s headline leaves out. In July 2023, Anthropic was one of seven companies — alongside Amazon, Google, Inflection, Meta, Microsoft, and OpenAI — that signed the White House’s voluntary AI commitments, which included developing marking mechanisms for AI-generated audio and image content. Text wasn’t on that list: it’s the hardest format to watermark without degrading quality, and the easiest to strip with a rewrite.
OpenAI, meanwhile, had already solved the technical problem years earlier. According to internal documents cited by the Wall Street Journal, the company built a watermarking tool for ChatGPT that was 99.9% effective on sufficiently long text, and sat on it for nearly a year without shipping it. The reasons it gave were, in themselves, a confession: nearly 30% of surveyed users said they’d use the product less if they knew their text would be marked, and the company worried the mark would disproportionately stigmatize people who don’t write in English as a first language, whose style already triggered more false positives in older detectors. The public classifier OpenAI did ship, in January 2023, was quietly retired six months later over its poor accuracy.
Google DeepMind, for its part, had been applying SynthID to Imagen’s images since 2023, extended it to text in 2024, and by May 2026 had watermarked more than ten billion pieces of content across its products. The technology Anthropic is announcing as its own in August 2026 is, in fact, built on the same method DeepMind published two years earlier. No technical breakthrough unlocked text watermarking this week. What arrived was, simply, a deadline.
The calendar was set by a law, not a conviction
That deadline is August 2nd, 2026, when Article 50 of the EU’s AI Act took effect, requiring any provider of generative systems to mark its outputs — text included — in a machine-readable, detectable way. The Code of Practice that implements that article, drafted with input from the very companies it governs, calls for a layered approach: embedded metadata, an imperceptible watermark, and logging of generation events. Anthropic had been committed to that code for months; the August 11th announcement is, in large part, the execution of that signature.
What’s telling is what happens outside Europe. Anthropic chose to switch on watermarking worldwide, not just where the law demands it, which reads like a gesture of global responsibility until it’s set against what China has required since September 2025: explicit labels, visible to the user, and implicit labels, embedded as metadata, on every kind of synthetic content distributed on platforms like WeChat, Douyin, or Weibo, under threat of fines and service suspension. Against that hard obligation, the EU imposes a technical requirement without a generalized visible label, and the United States imposes none at all — only California’s AI Transparency Act is starting to require something similar for audio, image, and video. The map that results is the usual one: a handful of regulators mandate something, and every other provider adopts the lightest version of that mandate as a de facto standard, presenting it as its own initiative.
A mark serves, above all, whoever issues it
It’s worth asking what Anthropic gains from this, beyond complying with a law that bound it regardless. The answer has two layers. The first is legal: facing a regulator, a lawsuit over the spread of disinformation, or a journalistic investigation into synthetic content, the company no longer has to answer “we don’t know” — it can answer “we published a detection API,” whether or not there’s a practical way for anyone outside Anthropic to use it with any confidence. The second is competitive: in a market saturated with what the press started calling AI slop — generic, mass-produced content eroding trust in anything that smells of AI — being the company that “does something” about it is, on its own, a selling point to advertisers and platforms worried about brand safety.
Neither layer requires the mark to actually work well for the reader. Anthropic hasn’t published false-positive rates, hasn’t explained who will get access to the detection API beyond “plans to release it,” and hasn’t described any process for disputing a detection that harms someone. It’s the same absence we ran into this same month, when we looked at who audits the self-reported safety evaluations AI labs publish: the company designs the test, the company decides what counts as a positive, and the company decides, later, who gets to verify that result.
The cost of doubt lands downstream
A mark that doesn’t distinguish well is still going to be used as if it distinguished perfectly, and Anthropic isn’t the one who’ll pay for that error. This has already happened before, without any watermarking involved: the previous generation’s style-based detectors — Turnitin, GPTZero, and the like — have spent two years producing false positives that have disproportionately hurt students who don’t write in English as a first language and authors with unconventional style, with no formal appeals mechanism in place. In August 2026, two concrete cases show where this same problem is headed with a supposedly more precise mark: literary agent support for Jerry Falade’s novel “Call Me, I’ll Hide the Body” was withdrawn over AI-use allegations the author denies, and Hachette pulled the horror novel “Shy Girl” after its author, Mia Ballard, told the New York Times she hadn’t used AI and suspected a freelance editor had introduced the flagged material without her knowledge.
A mark that says “Claude had a hand in this” doesn’t distinguish between those two scenarios and someone who mass-generated a thousand filler articles to monetize ads. But a school, a publisher, or a platform receiving that mark isn’t going to treat it as the imprecise volume filter it actually is — it’s going to treat it as proof, because treating it as proof is cheaper than investigating case by case. The mark’s technical design — deliberately weak, easy to strip with a full rewrite, unable to say how much of a text is human in origin — shifts Anthropic’s own ambiguity onto whoever receives that mark without the standing to question it.
Marking is not answering for what’s marked
None of this makes watermarking a bad idea. A world with imperfect watermarks is better than one with none, and the alternative — style-based detectors trained to guess, with no statistical guarantee behind them — already proved worse. The problem isn’t that Anthropic built this technology. It’s the gap between what the mark can actually do and what it will be expected to do, a gap the company itself acknowledges in the fine print of its own announcement and that never makes it into the headline.
Watermarking wasn’t designed, first and foremost, so the reader would know. It was designed so the company could show, to whoever asked, that it tried. That difference explains why the mark is deliberately weak, why it took three years to reach text after being promised for audio and image, and why nobody has yet answered the more uncomfortable question: what a school, a publisher, or a reader is supposed to do when the only evidence they have is a mark that its own maker warns shouldn’t be taken as evidence.
Sources
- Anthropic says it will watermark text generated by its AI models (TechCrunch)(techcrunch.com)
- Anthropic shares more details about how Claude's new watermarks will work (TechCrunch)(techcrunch.com)
- Anthropic Reveals What The Watermark Is And How It Can Be Defeated (Search Engine Journal)(www.searchenginejournal.com)
- Anthropic plans to add an invisible mark to AI text—as the industry scrambles to police AI slop (Fortune)(fortune.com)
- Claude Invisible Watermarks — What They Detect (And Miss) (explainx.ai)(www.explainx.ai)
- Watermarking AI-generated text and video with SynthID (Google DeepMind)(deepmind.google)
- The EU AI Act's Transparency Rules: A Practical Guide to Article 50(artificialintelligenceact.eu)
- China's AI Content Labeling Rules: Key Requirements & Actions (CMS Law)(cms.law)
- OpenAI has built a text watermarking method to detect ChatGPT-written content (Tom's Hardware)(www.tomshardware.com)
- FACT SHEET: Biden-Harris Administration Secures Voluntary Commitments from Leading Artificial Intelligence Companies(www.presidency.ucsb.edu)