The Hedge Does Not Travel With the Number

Anthropic now watermarks the text its models write. Peer-reviewed work shows the mark can be forged onto anyone's writing for about fifty dollars, and that strong watermarking is formally impossible. No error rate has been published. A legal, technical and philosophical examination of a detector aimed at the customer who paid for it.

Anthropic now watermarks the text its models write, and says the result proves only that Claude was “likely involved.” It has not published an error rate, a threshold, or a way to contest a result. The hedging protects the company’s statement. It does not travel with the number to the person who gets accused.

This piece reflects the public record as of September 2, 2026. Watermarking of Claude text began August 2, 2026 for models launched on or after that date, with older models described as in progress. The detection API has been described as a limited preview for qualified organizations. As of this writing, Anthropic has not published detection thresholds, false-positive rates, or a dispute procedure. Nothing here is legal advice.

Start with what was actually built, because almost every account of it is wrong in the same way.

When people hear “watermarked text,” they picture something inserted. Invisible Unicode. Zero-width spaces. A payload riding along in the file, waiting to be found. That is what most detection schemes have done, and it is what the phrase invites you to imagine.

That is not this. Nothing is added to the text and there are no hidden characters. When the model is choosing among words that are all equally good in context, the randomness behind that choice is seeded from a key. The words stay ordinary words. Afterward, someone holding the key can examine the sequence of choices and ask whether it matches what that key would have produced. The watermark is not in the text. The watermark is the text.

Two consequences follow immediately, and they cut in opposite directions.

The first is that this survives things metadata never survives. Print the page and scan it back, and optical character recognition returns the same words in the same order, which is the same fingerprint. Screenshot it, retype it character for character, move it between formats, strip every byte of metadata. None of it matters, because there was never anything in the bytes. That is genuinely clever, and it is the strongest thing about the design.

The second is that it dies to a five-minute effort. Anything that changes the words destroys the signal. Paraphrase it. Run it through a different model. Translate it and translate it back. Anthropic says as much: heavily edited, paraphrased, translated, or mixed writing loses the mark. Short passages carry too little signal to read at all.

So the system reliably identifies the person who pasted the output straight through without reading it, and reliably misses anyone who spent ninety seconds hiding it. Hold that thought, because it determines who actually gets caught by this, and the answer is not the people you would want caught.

What the company got right, stated fairly

This is a real regulatory obligation, not a marketing exercise. Article 50 of the EU AI Act requires providers to mark synthetic content in machine-readable form. Anthropic applied the marking worldwide rather than carving out a compliant region, which is more than the law demanded.

The engineering choice is defensible on its own terms. Sampling-based watermarking does not degrade the writing, does not push toward stilted synonyms, and does not litter documents with characters that break when they meet a PDF pipeline. For images, the company attaches signed C2PA provenance metadata, which is an open standard rather than a proprietary one.

And the public description is careful. A positive result does not identify authorship or ownership. It does not prove a terms-of-service violation, fraud, or the truth of anything. A negative result does not prove a human wrote it. Those are honest sentences, and they are more restrained than most vendors manage.

The problem is not that Anthropic overclaimed. The problem is what happens to a carefully qualified number after it leaves the building.

The defect is the missing denominator

A detection system makes an accusation-shaped statement about a specific person. The single most important fact about any such system is how often it is wrong, and in which direction.

That number has not been published. Not the false-positive rate. Not the thresholds. Not the calibration. Not the confidence intervals, and not the minimum length below which the tool should refuse to answer at all. There is also no disclosed dispute process, no appeal, and no stated way for an accused person to obtain the underlying computation.

Anthropic has this data. You cannot ship a detector without measuring it. The decision was to release the capability and withhold the measurement.

That asymmetry is the whole issue. The institution gets a number. The person the number is about gets nothing: no rate, no method, no recourse, and no standing to demand any of it.

The false positive that matters is not the one people imagine

Give the mathematics its due. This is not the perplexity-and-burstiness guesswork that has been flagging non-native English speakers as cheaters for three years. A key-based test is a real statistical test, and the probability that a person independently produced a long passage matching a key’s choice pattern falls away exponentially with length. Someone who never touched the tool being flagged from nowhere is not the realistic failure.

Here is the realistic failure. Detection does not distinguish between the model having written the text and the model having edited it.

Sit with what that means. A person writes something entirely their own. Their argument, their evidence, their voice, months of it. They run it through Claude for a grammar and punctuation pass, the way three generations have run work through a spell checker. Enough of the resulting word choices are the model’s that the passage carries signal. The check comes back positive. The output does not say “lightly edited,” because it cannot say that. It says the model was involved.

The work is theirs. The result reads as machine. There is no field in the response that can tell the difference, and no threshold the person can point to, because none was published.

There is a second version that matters for anyone who publishes original work. A third party takes your writing, feeds it to the model, and asks for a reproduction or a light rewrite. The output carries a mark. Now a marked near-copy of your language is circulating. The mark attaches to the copy rather than to you, but it has put a machine-flagged version of your words into the world, and you did not do it and will not know.

The attack is published, it costs fifty dollars, and it runs backwards

Everything above assumed the mark is at least honest when it is present. The peer-reviewed literature does not support that assumption, and this is the part of the record that should end the conversation about whether any of this is ready to be used against a person.

At ICML 2024, researchers at ETH Zurich published Watermark Stealing in Large Language Models. An attacker holding nothing but ordinary public API access can approximate the provider’s secret watermark rules for under fifty dollars in queries. With that approximation, the attacker can forge the watermark onto text of their own choosing, at better than eighty percent success while keeping the writing high quality. The same approximation works in reverse, lifting removal success against one scheme from roughly one percent to over eighty. The authors state their conclusion without hedging: current watermarking schemes are not ready for deployment.

Now read that against the accusation the detector produces.

If the mark can be manufactured for the price of a dinner, then a positive result does not mean the model was involved. It means the model was involved or somebody wanted it to look that way, and the detector cannot distinguish those two cases. There is no signature, no key in the hands of the accused, nothing to check the claim against. The output is a bare assertion that the accused has no instrument to rebut.

Consider who that arms. A competitor for a contract. An opposing party who would like an expert’s report to look fabricated. Somebody who wants a rival’s manuscript, or a student’s thesis, or a pro se litigant’s brief to come back flagged. Every one of those attacks is now described in a published paper with a measured success rate, and the target has no way to demonstrate that the mark was planted rather than earned.

The second result is more fundamental. Also at ICML 2024, Watermarks in the Sand proved that strong watermarking is impossible under natural assumptions: a computationally bounded attacker can erase the mark without significant quality degradation. The impossibility holds even in the private-detection setting, where insertion and detection share a secret key the attacker never sees. The authors demonstrated the attack against three published schemes, removing every watermark with only minor quality loss.

Put the two together and the standing academic position is that the mark is removable as a matter of proof and forgeable as a matter of practice, at a price anyone can pay.

A system with those properties has inverted its own purpose. It does not catch the sophisticated, who can strip it. It does not protect the innocent, who can be framed with it. What it reliably identifies is the careless, and what it reliably supplies to the motivated is a weapon. Deploying that against people, while publishing no error rate, is not a close call.

You cannot disclaim a duty to a stranger

The most common defense of all this is that the qualifying language solves it. It does not, for a reason that has nothing to do with how well the qualifiers are written.

A disclaimer runs to the party who receives it. The person who gets flagged is not that party. They never contracted with the provider, never saw the terms, never clicked anything, and never had the opportunity to accept or refuse. There is no privity. You cannot allocate away a duty owed to someone who was never at the table, and no amount of careful drafting in an agreement with a university changes what is owed to the student that university then accuses.

The qualifiers do real work protecting the provider’s own statement from being called false. They do nothing about a claim brought by the third party the statement was about.

And the qualifiers do not travel. What reaches the accused person is not “probabilistic signal consistent with possible involvement, error rate unpublished.” What reaches them is “our system says this is AI-generated.” The hedge stays in the documentation. The accusation is what arrives.

Foreseeable use is not an intervening cause. It is the market.

The next defense is that any harm is caused by whoever misuses the score, not by the company that produced it.

An intervening act cuts off proximate cause when it is unforeseeable. This one is the business model. A detection API sold to institutions is sold for exactly one purpose: deciding things about people. Universities adjudicating misconduct, employers screening submissions, publishers evaluating manuscripts, agencies reviewing proposals. Nobody is buying this to satisfy idle curiosity.

The Restatement puts it directly. Where the likelihood that a third person will act in a particular way is the very hazard that makes the actor’s conduct negligent, that act does not relieve the actor of liability. Selling an unvalidated accusation engine into the accusation market is not undone by the fact that the customer pulled the trigger.

The doctrine that fits is not defamation

Defamation is the theory everyone reaches for, and it is the weakest one available here. Hedged probabilistic language, disclaimed authorship, no assertion of misconduct. That is difficult ground, and the provider’s caution was almost certainly drafted with it in mind.

Though it is not the shield it appears to be either. Milkovich v. Lorain Journal Co. settled that labeling something an opinion is not a talisman. A statement that implies the existence of undisclosed facts supporting it remains actionable. “Our analysis indicates this was machine-written” implies a validated method behind it. When the method’s error rate is deliberately unpublished, the implication is doing work the disclosure does not support.

But the cleaner fit is negligent misrepresentation, Restatement (Second) of Torts § 552. One who, in the course of business, supplies false information for the guidance of others in their transactions is liable to those the information is intended to influence, for losses caused by justifiable reliance, where reasonable care was not exercised in obtaining or communicating it.

That formulation does not require a defamatory statement of fact. It asks whether an information supplier took reasonable care in how the information was communicated. Releasing a scoring system with no published error rate, no minimum-length refusal, no authorship-versus-editing distinction, and no dispute channel, into a market known to use it punitively, is a failure-of-care case on its face. The careful qualifiers are not a defense to it. A plaintiff would offer them as evidence the company understood the limits and shipped anyway.

The product framing runs parallel. A design-defect claim needs a reasonable alternative design, and here there is an obvious one that costs almost nothing: publish the thresholds, refuse to score passages below a reliable length, state plainly that authorship cannot be separated from editing, and provide a route to contest a result. Every element is feasible. None was done. That is what a plaintiff’s expert puts on a board.

The evidentiary problem, and the venues where it will not save anyone

In a courtroom this evidence is weak. Rule 702 and the Daubert factors make the known or potential rate of error an explicit reliability criterion. A proponent who cannot state the error rate, because the developer has not published one, is in trouble before the merits. Competent counsel takes it apart on voir dire.

That protection is mostly theoretical, because almost none of the proceedings where this gets used are courtrooms. A university honor board. A journal editor. A grant panel. A contracting officer. A client who decides quietly not to call back. No rules of evidence, no cross-examination, no discovery, no reasoned decision, and frequently no appeal. The accused is handed a number they cannot inspect and asked to prove a negative.

A federal agency applied this exact test, this year

The standard being asked for here is not exotic or novel. A federal regulator applied it in 2026, to a different detector, and published its reasoning.

Section 24220 of the Infrastructure Investment and Jobs Act directs the National Highway Traffic Safety Administration to require passive impaired-driving detection technology in new passenger vehicles. In its March 2026 report to Congress, the agency effectively declined to mandate it, on the ground that available systems carry unacceptable error rates. Even at 99.9 percent accuracy, deployment at the scale of the American vehicle fleet produces millions of false positives a year.

That is the argument of this piece, made by an agency, about a detector, in public, with the arithmetic attached. NHTSA did not claim the technology was worthless or that the underlying goal was illegitimate. It said the error rate at deployment scale was unacceptable, and it showed the work.

So one detector must clear a published error threshold before it is permitted to stop a car. Another is deployed into decisions about degrees, contracts, publication and professional reputation with no published rate whatsoever. Both are detectors. Both are wrong sometimes. Only one had to say how often.

One structural point cuts in favor of future plaintiffs. Section 230 does not apply. This is not third-party content a platform hosted. It is the provider’s own output, about a person. The immunity that has absorbed most technology liability for a quarter century is simply not in the room.

The regulation that created it may be the one that reaches it

There is an irony in the sequence worth naming.

The marking exists to satisfy Article 50 of the EU AI Act. But once a detection score drives a consequential decision about a person, other law engages. GDPR Article 22 restricts decisions based solely on automated processing that produce legal or similarly significant effects. Article 15(1)(h) gives the data subject a right to meaningful information about the logic involved.

A provider that has published no thresholds, no error rate, and no dispute path has a difficult time answering an Article 15 request, and the institutions relying on the score inherit the exposure. The compliance measure generates the compliance problem.

The tobacco parallel, stated precisely

The analogy is invoked loosely and it deserves better, because the specific mechanism transfers and the general moral does not.

The tobacco industry did not fall because cigarettes cause cancer. That was known for decades while the industry kept winning. It fell on failure to warn and fraudulent concealment, and it fell in discovery, when internal research surfaced showing what the companies had measured and chosen not to disclose. The gap between private knowledge and public statement was the case.

The structure maps. A company possesses internal measurements of how often its product harms people. It publishes a carefully worded public account that does not include them. If those measurements are materially worse than the public framing suggests, that gap is the entire litigation, and it is discoverable.

Where the analogy should be tempered: tobacco took forty years, individual plaintiffs lost for most of them, and the vehicle that finally worked was state attorneys general suing to recover Medicaid costs, with subpoena power and institutional patience no private plaintiff has. The realistic sequence here looks similar. Consumer-protection and unfair-practices enforcement by a state attorney general, or a European regulator, will reach the internal numbers long before a lone plaintiff does. Early private suits will founder on damages and causation.

Marketing a detection service while withholding its error rate is a cleaner deceptive-practices story than it is a tort story. And discovery on the withheld measurement is not a side issue in that case. It is the case.

Nobody watermarks a ghostwriter

Step back from the liability question and ask the one underneath it, because the answer is less obvious than the whole debate assumes.

Ghostwriting is lawful, ancient, and entirely unmarked. Under the Copyright Act, work prepared by an employee within the scope of employment, or commissioned under a signed work-for-hire agreement in one of the enumerated categories, belongs to the hiring party, who is treated as the author as a matter of law. That is not a loophole. It is the statute.

Presidential addresses. Memoirs. Op-eds under executives’ bylines. Nearly all corporate communication. Every one of those is words produced by one person and published under another person’s name, and none of it is disclosed, marked, detected, or scandalous. A reader of a chief executive’s op-ed is exactly as uninformed about who composed the sentences as a reader of a machine-assisted draft.

So the distinction being drawn is not about what the reader knows. Pay a human being to write in your voice and you are the author. Use a tool that writes in your voice and the output carries a mark for the life of the document.

Every other trade settled this long ago. A tailor works from a pattern and a computer-guided cutter. A chef uses a food processor and stock made by somebody else. A photographer uses autofocus, and before that a light meter, and before that a lens somebody else ground. In none of those fields does the output carry a label describing which tools touched it, and in none of them does anyone experience this as a scandal waiting to be uncovered.

Construction is the most instructive case, because construction is saturated with mandatory disclosure and none of it is aimed at authorship. Stamped drawings, sealed calculations, permits, inspections, code compliance, lien filings. Not one page records which carpenter drove which nail, or whether the trusses were cut on site or arrived prefabricated. What the entire apparatus records is who is accountable. A licensed engineer stamps the drawing and owns what happens next.

That is the mature form of this problem. It was worked out in a field where the failure mode is a collapsed building, and it landed on accountability rather than provenance of effort. Nobody needed to detect the engineer. The engineer signed.

Authorship is younger than it feels

Underneath the detection impulse sits an intuition: that a text has one true human origin, that this origin is a real property of the words, and that the right instrument could recover it. That intuition feels ancient. It is roughly three hundred years old, and it was built for a commercial purpose.

Foucault’s What Is an Author? made the case that the author is not a natural fact about a text but an author-function: a way a society classifies discourse, limits it, and above all decides who answers for it. Barthes had already argued that meaning is assembled by the reader rather than transmitted from an origin. Neither was making a clever academic point. Both were describing something the law had already discovered, which is that attribution is an assignment of responsibility.

The historical work is more pointed still. Martha Woodmansee showed that the modern figure of the author as solitary originating genius emerged alongside the eighteenth-century book trade, which needed a stable owner for a new kind of property. Peter Jaszi traced how that romantic construct then shaped, and distorted, copyright doctrine itself. The idea that creation is a single mind producing something from nothing was not observed in nature. It was useful, and it was adopted.

Which is why the law never actually implemented it. Work for hire, joint authorship, editorial revision, and the entire apparatus of professional signature all point the same direction: the system tracks who is responsible, not who had the thought. Detection is an attempt to build the metaphysics the doctrine deliberately declined to build.

Walter Benjamin described what happens when reproduction technology arrives: the aura attached to the singular object dissolves, and the culture relocates value rather than losing it. Photography did not end painting. It ended the belief that manual reproduction was where the value lived. And on the account Andy Clark and David Chalmers gave of the extended mind, a tool that participates in the thinking is part of the cognitive system rather than a contaminant of it, which is the same reason nobody has ever argued that an outline, a thesaurus, or a research assistant makes a person less the author of their own argument.

None of this says nothing has changed. It says the thing being defended, a single detectable human origin, is a recent invention that the law itself never adopted, and that we are now proposing to enforce it by forensics for the first time in its history, at the exact moment the forensics are proven unreliable.

The exhaustion is not hypothetical

There is also a predictable end state to marking everything, and we have run this experiment already.

California’s Proposition 65 required warnings for a growing list of substances until the warnings appeared on parking garages, coffee, and hardware stores. The label is now ambient. It conveys nothing, because it is attached to everything, and the people it was written to protect walk past it without reading it. Cookie consent banners produced the same result on a shorter timeline: universal, ignored, and clicked away reflexively by people who wanted only to reach the page.

Over-disclosure does not add information. It destroys the channel the information was supposed to travel through. A marking regime broad enough to cover every assisted document will arrive at exactly that place, and the arrival will be quiet, because nobody announces the moment a warning stopped meaning anything.

Where it does matter, stated narrowly

The argument above is not that nothing should ever be disclosed. Three categories are genuinely different, and they share a structure worth naming.

Where a person certifies personal knowledge or judgment. Sworn testimony. A physician’s diagnosis. An auditor’s opinion. An attorney’s signature under Rule 11, which certifies that the filing is grounded in fact and warranted by law after a reasonable inquiry. In each, the value of the document is the human judgment behind it, and a person who did not exercise that judgment has misrepresented something regardless of what tool was involved.

Where the artifact is being used to measure the person rather than to do a job. A graded examination is not trying to obtain a good essay. It is trying to find out what the student can do. That is a genuine and defensible distinction, and it is the one honest case for detection in education. It is also much narrower than how detection is actually being deployed there.

Impersonation and fraud. Passing work off as a specific other person’s, or fabricating a record. Those were already wrongs and did not need a new category.

Outside those three, the honest answer to what is at stake is: very little. A finished thing is good or it is not. It is accurate or it is not. Someone stands behind it or nobody does.

And notice that the professional rules already work this way. The duty stays on the lawyer, the physician, the engineer, whatever assistance was used in getting there. Bar guidance on generative AI has kept the obligation exactly where it was, on the professional who signs. That architecture is correct, it is already built, and it does not require detecting anything.

Everything sits on something else

There is a claim buried under the whole detection project that nobody states out loud, because stating it would expose it. The claim is that human creation is categorically unlike machine creation, and that the difference is detectable in the output.

Look at how a person actually acquires the ability to write.

You did not invent your language. You absorbed it. Nobody sat you down and taught you the grammar before you started using it correctly; you were surrounded by enough of it that the pattern formed. Your accent is a map of where you lived. Your sentence rhythm is a residue of what you read. Every writer who ever got good got there by reading other writers and, at first, sounding like them. In every art form that has a pedagogy, copying the masters is the pedagogy, and nobody has ever called it cheating.

This is not a loose analogy. It is among the most replicated findings in the behavioral sciences.

In 1996 Saffran, Aslin and Newport put eight-month-old infants in front of a continuous stream of nonsense syllables with no pauses, no stress cues, and no meaning. After two minutes, the infants had extracted the word boundaries, using nothing but the transitional probabilities between sounds. Infants perform statistical learning over a corpus. That is the finding, in humans, before they can walk.

Albert Bandura had already established that children acquire entirely novel behaviors from observation alone, with no reinforcement, no instruction and no reward. Chartrand and Bargh later documented the chameleon effect: adults unconsciously adopt the postures, gestures and speech patterns of whoever they are with, without intending it and without noticing. The mimicry runs on its own, below the level of decision.

You are, in a real and measurable sense, an accumulation of your exposure. The language, the environment, the people, the culture, the music, the work you were drawn to. The mind copies what it attends to, whether or not the person is trying.

The neuroscience argument, where there is one, is about which circuitry does this, not about whether it happens. In the early 1990s Giacomo Rizzolatti’s group in Parma found neurons in the macaque premotor cortex that fired both when the animal performed an action and when it watched someone else perform it. A strong version of the claim followed, that these mirror neurons are the innate machinery of understanding other minds, and Gregory Hickok challenged that strong version in 2014.

But look where the challenge lands. The leading alternative, developed by Cecilia Heyes, Caroline Catmur and colleagues, holds that mirror neurons are not an evolutionary endowment at all: they are built after birth by associative learning, formed through repeated instances of doing and seeing the same action at the same time. So the two competing accounts are an innate imitation system, or an imitation system assembled from correlated exposure. Both roads arrive at learning by observation. The disagreement is over the wiring diagram, and it leaves the thing that actually matters here untouched.

And notice the direction the skeptical account points. If the mirroring capacity is itself a statistical residue of enough correlated experience, then the most defensible neuroscience of human imitation is describing training.

Zoom out and the pattern is the whole story of the species. Michael Tomasello’s work on cumulative cultural evolution describes the ratchet effect: an innovation is transmitted, retained, and improved by the next generation, which never has to rediscover it. No human being has ever built anything from nothing. Every sentence any of us writes is assembled from a language we did not invent, using forms we inherited, about ideas we encountered, in a genre somebody else established. Bernard of Chartres said it in the twelfth century and Newton repeated it: we see further because of what we are standing on.

The law understood this before the neuroscience did. Copyright protects expression and refuses to protect ideas, facts, methods and systems, because those are the common stock everything is built from. Feist requires only independent creation plus a modicum of creativity, a deliberately low bar that concedes almost everything is recombination. Scenes a faire and the merger doctrine exist precisely to keep the inherited and the inevitable out of anyone’s private ownership. The entire architecture is an admission that originality is a thin layer over an enormous shared substrate.

So we built a system that ingests the accumulated written output of the species, finds the patterns in it, and recombines them into new arrangements. Then we described that as categorically different from what a person does, and set out to detect the difference in the words.

The reasonable position is not that these are the same thing. They are not. A person has a body, a life, stakes, and something to lose. The reasonable position is narrower and harder to escape: the difference is not located in the text. It is located in who is answerable for it. Which is exactly where the professional signature has always put it, and exactly what a detector cannot see.

The slippery slope runs the direction nobody expects

And here is the part that should give everyone pause, whichever side of this they are on.

Law confers status by regulating. A thing that must be marked, disclosed, restricted and accounted for is a thing the legal system has begun to treat as an actor rather than an instrument. Nobody legislates the disclosure obligations of a hammer.

The doctrine is currently pointed the other way, and firmly. In Thaler v. Vidal the Federal Circuit construed the Patent Act’s definition of an inventor to require a natural person, excluding a machine. In Thaler v. Perlmutter the D.C. Circuit affirmed the parallel rule in copyright in March 2025, and the Supreme Court declined to review it in March 2026, so the human-authorship requirement is now final rather than merely persuasive. The Ninth Circuit’s monkey-selfie decision in Naruto v. Slater supplies the broader proposition that non-humans do not hold these rights. Machines cannot author, and that question is now closed at the top of both systems.

Which produces a genuine doctrinal tension worth naming plainly. We hold that machine output has no author, and simultaneously build a forensic apparatus premised on the idea that machine output bears a detectable authorial hand. It cannot be authorless for the purpose of ownership and authored for the purpose of accusation. One of those has to give.

And the historical warning is right there in the corporate form. The corporation began as a convenience, a way to hold property and be sued. Centuries of incremental regulation and protection later, it holds constitutional rights. Nobody designed that outcome. It accumulated, one reasonable increment at a time, because each new obligation implied a new capacity, and each new capacity eventually implied a right.

Every rule written to constrain these systems is also a rule that treats them as the kind of thing rules are written about. That is not an argument against regulating them. It is an argument for regulating the people and companies deploying them, which is where the accountability actually lives, rather than building an ever-larger body of law addressed to the artifact itself. The first approach keeps a person on the hook. The second, followed far enough, produces a legal subject.

The customer is the one being marked

There is an asymmetry here that people feel before they can name, and it is worth naming.

The person whose writing carries the mark is the paying customer. Not an intruder, not an adversary, not somebody who obtained the product improperly. Someone who bought a service, used it exactly as intended, and received work product carrying a property that can later be used to make an accusation against them, for the benefit of a third party who wants to check up on them.

No other purchased tool behaves this way. A word processor does not mark your documents so somebody can later determine you used a word processor. A camera does not brand the photograph so an editor can rule you did not paint it. A commissioned ghostwriter does not embed a signal in your memoir for a publisher to find. In every one of those relationships, the output belongs to the buyer and carries nothing that runs against the buyer’s interest.

When does it end, and the answer nobody likes

The natural next question is where this stops, and it is usually asked rhetorically, as though the absurd end state were far off. It is not far off. Most of it shipped while this was being argued about.

  • Advertising inside the paid word processor. Microsoft displays a persistent banner in the title bar of Word, Excel and PowerPoint prompting users to upgrade their plan. It is shown to people who already hold an active paid Microsoft 365 subscription, and there is no user control to remove it. Commercials in the document window, delivered to the customer who already paid.
  • Advertising on the refrigerator. Samsung began running banner advertisements on the screens of Family Hub refrigerators retailing above eighteen hundred dollars. Owners can switch the ads off, but doing so also disables the news, weather and calendar widget they bought the screen for. The choice offered is the advertisement or the feature.
  • Advertising in the subscription that existed to avoid advertising. Amazon moved Prime Video subscribers onto an ad-supported tier and charged an additional fee to opt back out of the commercials they had not previously been shown.
  • Identity as a precondition of access. In Free Speech Coalition v. Paxton (2025) the Supreme Court upheld a state age-verification requirement under intermediate scrutiny. Proving who you are before you are permitted to read something is no longer a hypothetical about the future. It is precedent.

Cory Doctorow named the pattern enshittification: a platform is good to its users, then degrades their experience to serve its business customers, then extracts from those customers too. The mechanism is not mysterious and it is not malice. It is simply that each increment is small, the customer has already committed, and switching is expensive.

What makes it durable is that the legal architecture is already built. Section 1201 of the Digital Millennium Copyright Act makes circumventing an access control unlawful even when the underlying act would be perfectly lawful, which is how printer cartridges, tractors and game consoles came to remain tethered to their manufacturers after the sale. Once a restriction is technically enforced, removing it becomes a legal offense rather than a consumer choice, and the right-to-repair fight of the last decade is the record of what that costs.

Watermarking is the same move in a new place. A capability is added to a product the customer already paid for. The customer cannot decline it, cannot inspect it, and cannot remove it without stepping outside the terms. The benefit runs to somebody else. And each increment is defended as modest, on the grounds that the previous increment already happened.

So the honest answer to when it ends is: it ends where customers and regulators make it end. It has never once ended because a company reached a limit on its own.

What this means if you write for a living

Practical, and available now.

Keep the composition trail. A probabilistic score loses to dated, incremental, human evidence of how a document came to exist. Version history. Timestamped commits. Drafts that grow across weeks. Research that predates the writing. That record survives editing, it is contemporaneous, and it answers the question the watermark only gestures at. Anyone working in version control already generates it and should stop treating it as a developer habit.

Understand that a proofreading pass is no longer evidentially free. Not that it is wrong, and not that it should be avoided. But for anything where authorship may be adjudicated later, a grammar pass now leaves a trace that cannot be distinguished from authorship, and you should know that before rather than after.

Prefer provenance you control. Detection by inference asks a stranger to guess who wrote something. Signing asks nobody to guess. A cryptographic signature over a specific artifact at a specific moment, anchored so a third party can verify it independently, attests to a state rather than estimating an origin. It survives editing, because it is not trying to read authorship out of word choice. It tells you what changed and when, which is the question people actually care about.

That is the distinction the whole debate keeps missing. Watermarking is a guess about origin. Provenance is a record of state. The first degrades the moment anyone touches the work. The second is the only one that holds up when something is contested, and it is the one nearly nobody is building.

The principle, stated plainly

A company may publish an accusation-shaped number about a person who never agreed to anything, while withholding how often that number is wrong, and disclaim the consequences to a party who never received the disclaimer.

That arrangement does not survive contact with ordinary tort law, and it should not. Not because the engineering is bad, because it is not. Because measuring the harm, knowing the measurement, and publishing everything except the measurement is a pattern the law has seen before and has already decided how to treat.

But the deeper error is upstream of the liability. We are building detection for a question the working world settled generations ago, and settled correctly. Every serious trade regulates accountability, not authorship. The engineer stamps the drawing. The auditor signs the opinion. The attorney signs the filing. None of those systems ask who held the pencil, and all of them work, because the signature is what carries the weight.

Detection asks a stranger to guess how a thing came to exist. Accountability asks a named person to stand behind it. One of those degrades the moment somebody edits a sentence. The other has held up buildings for a century.

The tell is not the watermark. The tell is the missing denominator, and behind it, the wrong question.


Sources and Authorities

I. The watermarking program (primary sources):

  • Anthropic, “How Claude’s text watermarking works,” https://www.anthropic.com/news/claude-text-watermark. The technical account: the mark changes the source of randomness in selection among equally valid word choices; “Nothing is added to the text and there are no hidden characters”; light revision likely preserves the signal while a complete rewrite eliminates it; minimal watermarking where Claude proofreads human text; detection API in limited preview.
  • Anthropic Help Center, “How Claude marks AI-generated content,” https://support.claude.com/en/articles/16266773-how-claude-marks-ai-generated-content. Scope across the API, Claude, Claude Code, Claude Cowork and Claude Tag, plus AWS, Google Cloud and Microsoft Foundry; marking supported on models launched on or after August 2, 2026, with older models in progress; C2PA signed provenance metadata for supported file types including .svg, .png and .jpg; signal may be lost where text is heavily edited, paraphrased, translated, or mixed into other writing.

II. Stated limits of detection:

  • XenoSpectrum, “What Can Claude’s Text Watermark Actually Prove,” https://xenospectrum.com/en/anthropic-claude-text-watermark-global-rollout-limits/. A positive result does not identify authorship or ownership and does not prove a terms violation, fraud, or the truthfulness of content; a negative result does not prove human authorship; detection does not distinguish whether the model wrote or substantially edited the text; no disclosed release date, pricing, eligibility, or dispute mechanism.
  • Axios, “Anthropic’s text watermarks signal new front in AI detection” (Aug. 12, 2026), https://www.axios.com/2026/08/12/anthropic-claude-watermarks-ai-detection.

III. Regulatory framework:

  • Regulation (EU) 2024/1689 (Artificial Intelligence Act), Article 50, transparency obligations including machine-readable marking of synthetic content.
  • Regulation (EU) 2016/679 (GDPR), Article 22 (decisions based solely on automated processing producing legal or similarly significant effects) and Article 15(1)(h) (right to meaningful information about the logic involved).

IV. Tort and product authorities:

  • Restatement (Second) of Torts § 552, Information Negligently Supplied for the Guidance of Others.
  • Restatement (Second) of Torts § 449, intervening acts that are the very hazard making the conduct negligent.
  • Restatement (Second) of Torts § 578, liability of a republisher.
  • Restatement (Third) of Torts: Products Liability § 2(b) (design defect, reasonable alternative design) and § 2(c) (inadequate warning).
  • Milkovich v. Lorain Journal Co., 497 U.S. 1 (1990). No wholesale exemption for anything labeled opinion; a statement implying undisclosed defamatory facts remains actionable.

V. Authorship, accountability, and disclosure fatigue:

  • 17 U.S.C. § 101, definition of a “work made for hire,” and 17 U.S.C. § 201(b), vesting authorship in the employer or commissioning party. The statutory basis for uncredited and undisclosed ghostwriting.
  • Fed. R. Civ. P. 11(b), certification by signature that filings are grounded in fact and warranted by existing law after an inquiry reasonable under the circumstances. Accountability attaches to the signer, not to the drafting method.
  • ABA Standing Committee on Ethics and Professional Responsibility, Formal Opinion 512 (July 29, 2024), “Generative Artificial Intelligence Tools,” retaining the obligations of competence, confidentiality, communication, candor, supervision and reasonable fees on the lawyer rather than the tool.
  • Cal. Health & Safety Code § 25249.6 (Proposition 65) and its warning regime, the standard example of a disclosure requirement applied so broadly that the warning ceased to carry information.

VI. Watermark robustness (peer-reviewed):

  • Nikola Jovanović, Robin Staab & Martin Vechev, Watermark Stealing in Large Language Models, ICML 2024 (SRI Lab, ETH Zurich), https://watermark-stealing.org/. Approximation of the secret watermark rules from public API access for under $50; spoofing (forging the mark onto attacker-chosen text) at over 80% success; scrubbing improved from roughly 1% to over 80% against KGW2-SelfHash. Authors’ conclusion: “Current watermarking schemes are not ready for deployment.”
  • Hanlin Zhang, Benjamin L. Edelman, Danilo Francati, Daniele Venturi, Giuseppe Ateniese & Boaz Barak, Watermarks in the Sand: Impossibility of Strong Watermarking for Generative Models, ICML 2024, https://arxiv.org/abs/2311.04378. Proof that strong watermarking is impossible under natural assumptions; the result holds in the private-detection setting where insertion and detection share a secret key unknown to the attacker; demonstrated removal against three published schemes with only minor quality degradation.

VII. Detection error rates in another regulated setting:

  • Infrastructure Investment and Jobs Act, Pub. L. No. 117-58, § 24220, directing a federal motor vehicle safety standard for advanced impaired-driving prevention technology.
  • NHTSA, Report to Congress: Advanced Impaired Driving Prevention Technology (Mar. 2026), https://www.nhtsa.gov/sites/nhtsa.gov/files/2026-03/Report-to-Congress-Advanced-Impaired-Driving-Prevention-Technology.pdf. Agency assessment that available detection technology carries unacceptable error rates, with even 99.9% accuracy yielding millions of false positives annually at fleet scale.

VIII. Consumer restriction, tethering, and identity as a precondition:

  • Cory Doctorow on enshittification, the sequence by which platforms move value from users to business customers and then to themselves.
  • 17 U.S.C. § 1201 (Digital Millennium Copyright Act), prohibiting circumvention of technological access controls irrespective of the lawfulness of the underlying use; the mechanism behind printer-cartridge, vehicle and console tethering, and the subject of the right-to-repair exemption proceedings.
  • Reporting on advertising placed into paid products: persistent upgrade banners in the title bar of Microsoft Word, Excel and PowerPoint shown to active Microsoft 365 subscribers; Samsung Family Hub refrigerator advertising pilots on units above $1,800, where disabling ads also disables the purchased widget; and Amazon’s move of Prime Video subscribers to an ad-supported tier with a surcharge to opt out.
  • Free Speech Coalition, Inc. v. Paxton, 605 U.S. ___ (2025), upholding a state age-verification requirement under intermediate scrutiny.

IX. Authorship theory:

  • Michel Foucault, What Is an Author? (1969), on the author-function as a historically contingent means of classifying discourse and assigning responsibility.
  • Roland Barthes, The Death of the Author (1967).
  • Martha Woodmansee, The Genius and the Copyright: Economic and Legal Conditions of the Emergence of the “Author” (1984), on the construction of romantic authorship alongside the modern book trade.
  • Peter Jaszi, Toward a Theory of Copyright: The Metamorphoses of Authorship (1991), on the influence of that construct upon copyright doctrine.
  • Walter Benjamin, The Work of Art in the Age of Mechanical Reproduction (1935), on the dissolution of aura and the relocation of value.
  • Andy Clark & David Chalmers, The Extended Mind (1998), on tools that participate in cognition as constituents of the cognitive system.

X. Learning, cumulative knowledge, and legal status:

  • Jenny R. Saffran, Richard N. Aslin & Elissa L. Newport, Statistical Learning by 8-Month-Old Infants, 274 Science 1926 (1996). Infants segment words from a continuous speech stream after roughly two minutes of exposure, using transitional probabilities alone.
  • Albert Bandura, social learning theory and the Bobo doll studies (1961, 1963), establishing acquisition of novel behavior through observation without reinforcement.
  • Tanya L. Chartrand & John A. Bargh, The Chameleon Effect: The Perception-Behavior Link and Social Interaction, 76 J. Personality & Soc. Psych. 893 (1999), on nonconscious mimicry of postures, mannerisms and expressions.
  • Giacomo Rizzolatti et al., the Parma findings on premotor neurons in the macaque firing during both execution and observation of an action.
  • Gregory Hickok, The Myth of Mirror Neurons: The Real Neuroscience of Communication and Cognition (2014), challenging the strong action-understanding hypothesis. The dispute concerns the neural mechanism, not whether observational learning occurs.
  • Cecilia Heyes, Caroline Catmur et al., the associative-learning account, under which mirror neurons are not an evolutionary adaptation but are formed after birth through repeated correlated experience of doing and seeing the same action.
  • Michael Tomasello, work on cumulative cultural evolution and the ratchet effect, by which innovations are transmitted, retained and improved across generations.
  • 17 U.S.C. § 102(b), excluding ideas, procedures, processes, systems and methods of operation from copyright protection; Feist Publications, Inc. v. Rural Telephone Service Co., 499 U.S. 340 (1991), on originality as independent creation plus a minimal degree of creativity, and the non-protectability of facts; and the scenes a faire and merger doctrines.
  • Thaler v. Vidal, 43 F.4th 1207 (Fed. Cir. 2022), construing 35 U.S.C. § 100(f) to require that an inventor be a natural person.
  • Thaler v. Perlmutter, No. 23-5233 (D.C. Cir. Mar. 18, 2025), affirming the human-authorship requirement for copyright registration; review declined by the Supreme Court in March 2026.
  • Naruto v. Slater, 888 F.3d 418 (9th Cir. 2018).

XI. Evidentiary standard:

  • Fed. R. Evid. 702; Daubert v. Merrell Dow Pharmaceuticals, Inc., 509 U.S. 579 (1993), identifying the known or potential rate of error as a reliability factor.
  • 47 U.S.C. § 230, inapplicable to a provider’s own first-party output.

XII. The tobacco record (for the concealment structure, not the product analogy):

  • Master Settlement Agreement (Nov. 1998), between forty-six state attorneys general and the major manufacturers, and the underlying state Medicaid recovery actions.
  • Cipollone v. Liggett Group, Inc., 505 U.S. 504 (1992), on the preemptive scope of federal labeling requirements and the survival of fraudulent-misrepresentation and conspiracy claims.

Shawn Paul Cosner, J.D.
Sparked Technology Solutions, Inc.

Share LinkedIn X
Subscribe to STS Talks

New posts in your inbox.

White papers, lab notes, findings, and briefings from Sparked Technology. No spam. One click to unsubscribe.

Want to discuss this further?

If your agency or organization needs sovereign AI infrastructure built around the principles in this post, we'd like to hear from you.

Request a Briefing