Anthropic is adding invisible watermarks to text produced by supported Claude models, including Claude Code, but the policy could also mark writing that a person created and Claude merely edited, translated or summarized. Until the company publishes a detector and reliability data, the signal may prove that Claude handled a document without proving who wrote it.
A teacher examining an essay may soon face a new kind of evidence. The paper could contain a hidden marker associated with Claude. That marker might indicate that an AI model generated the prose. It might also indicate that the student wrote the paper, then asked Claude to correct grammar, translate a passage or tighten several sentences.
Those are different events. A watermark may not be able to tell them apart.
Anthropic’s new model-level watermarking policy brings the distinction into focus. Reporting by TechCrunch and Ars Technica this month says text processed by supported Claude models can carry an invisible mark. The system is expected to apply not only to conventional Claude responses but also to output from Claude Code, Anthropic’s coding assistant. The reported scope may include text that Claude proofreads, summarizes, translates or lightly edits, rather than only text composed from an empty prompt.
Anthropic has not yet released a public detector that outsiders can test across real documents, model versions and editing workflows. It has also not published enough information to establish how often the marker survives rewriting, how frequently it appears in unmarked text or how reliably it separates generation from assistance.
That leaves the central question unresolved: what exactly is the watermark supposed to prove?
A mark of contact, not necessarily authorship
A watermark is a hidden pattern inserted into generated content. In text, that pattern can involve small statistical preferences in word choice or token selection. A token is a unit a language model uses when producing text. It may be a whole word, part of a word or punctuation.
The basic idea is simple. If a model has several plausible ways to continue a sentence, it can favor particular groups of tokens according to a secret rule. A detector with access to that rule can look for the pattern across enough text. The individual choices should remain natural to a reader, while the aggregate pattern identifies the model that produced them.
That description does not settle the harder problem. A language model may be involved in a document without being its author in the ordinary sense.
Consider a lawyer who writes a client letter and asks Claude to translate it into Spanish. Claude generates the translated text, but the ideas, factual claims and professional judgment originated with the lawyer. Or consider a software engineer who writes a function, then asks Claude Code to explain it in comments. The comments may carry a watermark even though the code and its underlying logic came from the engineer.
A similar issue appears in proofreading. If Claude changes ten sentences in a 2,000 word report, does the resulting document count as AI-generated? A binary detector may return a binary answer. The workplace does not operate in binary.
The distinction matters because institutions often use labels as shortcuts. A school may prohibit uncredited AI writing. An employer may require disclosure of generated material. A publisher may want to know whether a submission was written by a person. None of those policies necessarily treats translation, copy editing and original composition as the same act.
A marker that says “Claude touched this text” could be useful. A marker interpreted as “Claude wrote this text” could be misleading.
The policy is clearer than the technology
Anthropic’s reported policy is an operational decision, not a published scientific result. The company is choosing to attach a provenance signal to model output as transparency requirements develop. It has not, based on the reporting available, demonstrated the detector’s performance under a full set of public tests.
That difference is important. A watermark proposal can describe an intended mechanism. A working provenance system must answer measurable questions.
How much text does detection require? Does accuracy change between 100 words and 1,000 words? Does a detector identify a watermark in a single paragraph, or only in a long document? What happens when content contains quotations, code, names, citations or specialized terminology? Can a user copy and paste the text without affecting the signal?
The editing question is even more consequential. If a person rewrites half of a watermarked passage, does the detector still identify Claude? If so, does it report the original model involvement accurately, or does it overstate the role Claude played? If a second model paraphrases the text, does the watermark disappear? If a human translates the passage, what remains?
These are not edge cases. They are normal document workflows.
Watermarks tend to face a tradeoff between robustness and invisibility. A stronger statistical pattern may survive more transformations, but it can make text less natural or easier to detect and remove. A weaker pattern may preserve fluency, but disappear after ordinary editing. The system must also avoid false positives, which occur when unwatermarked text is incorrectly classified as marked.
A detector can be technically impressive and still be unsuitable for high-stakes decisions. A 99 percent accuracy rate sounds strong until an institution checks 100,000 essays or employee reports. At that scale, even a small false positive rate can produce many wrongful accusations. The actual risk depends on the base rate, meaning how common genuinely watermarked text is in the population being tested.
Suppose only a small fraction of submitted essays contain Claude-generated text. A detector with imperfect specificity could flag more legitimate essays than prohibited ones. That is a standard statistical problem, not a peculiarity of AI. A hidden mark does not eliminate it.
Anthropic’s public evidence will need to report precision, recall and false positive rates under defined conditions. Precision measures how many flagged documents actually contain the target signal. Recall measures how many marked documents the detector finds. Neither number has much meaning without the sample composition, text length and transformations used in the test.
The company will also need to disclose whether the detector is available to independent researchers. A vendor-operated detector can be useful for product support, but institutions that use the result to discipline a student or reject a submission need a way to audit the process. Otherwise, the watermark becomes an assertion backed by an inaccessible instrument.
Why Anthropic has an incentive to watermark
The policy serves more than one purpose.
The obvious rationale is provenance. As governments and institutions ask AI companies to make generated content identifiable, a model-level watermark offers a direct response. It gives Anthropic a way to say that Claude output is not being released without any trace of origin.
There is also a competitive incentive. Watermarking can become part of an enterprise trust package. A company buying Claude for customer service, software development or internal communications may want records showing where text came from. An identifiable output stream could support compliance teams and corporate content controls.
That does not mean the system solves provenance broadly. It would identify, at best, text associated with supported Claude models and a functioning detector. It would not identify text produced by another model, text generated locally or text copied from a human source. It would not establish whether a user accepted the output unchanged. It would not capture AI assistance that leaves no machine-generated wording behind.
Still, a vendor-specific signal has commercial value. It can help Anthropic answer questions from regulators and buyers while encouraging customers to remain inside its ecosystem. If a company uses Anthropic’s detector, logging tools and policy controls together, the customer becomes more dependent on Anthropic’s definitions and infrastructure.
Claude Code adds another layer. Software development already contains multiple forms of machine assistance. A model may generate a complete function, suggest a small fix, write documentation or summarize a codebase. Those activities have different consequences for ownership, review and security. A watermark on comments or natural-language explanations could help track tool use, but it would not by itself show whether a developer inspected the code or whether the code introduced a vulnerability.
The business question is therefore not only whether a watermark works. It is who gets to interpret it.
The workplace will need a vocabulary more precise than “AI”
A document review system that returns one label is likely to create confusion. Organizations need categories that reflect how people actually use these tools.
One useful distinction is between generation and transformation. Generation means the model produces new substantive content, such as an analysis, a report section or a software function. Transformation means the model changes existing content through translation, proofreading, formatting or summarization.
Even that distinction has limits. A summary is not merely a cosmetic transformation. It selects what to include and what to omit. A rewrite can alter meaning. A grammar correction can change tone or legal interpretation. Human assistance and machine assistance are not interchangeable categories.
A better disclosure record might include the tool, the task and the extent of the change. “Claude translated a human-written draft” tells a reader more than “AI used.” “Claude generated the first draft, then an editor revised it” tells a different story. Neither statement requires pretending that authorship is a single measurable property.
Teachers face the most difficult version of the problem because assessment is often designed to infer what a student can do without assistance. If a student submits a watermarked essay, the mark may prompt a useful conversation. It should not automatically establish misconduct. The teacher would still need to examine drafts, revision history, citations, discussion of the argument and the student’s ability to explain the work.
Employers have similar obligations. A company may want employees to disclose generated text for legal, security or quality reasons. But a policy that treats proofreading as equivalent to ghostwriting could discourage harmless use and push assistance underground. Employees may avoid approved tools, use unapproved ones or remove visible signs of assistance. The control system would then lose the very transparency it was meant to create.
Publishers and news organizations have another concern. A watermark could help editors identify material that needs disclosure or additional fact checking. It cannot replace those checks. Human writing can be inaccurate, copied or fabricated. AI writing can be accurate and heavily edited. Origin is relevant, but it is not the same as quality.
Watermarks can be attacked, and ordinary editing may be enough
Any provenance system must be evaluated against deliberate evasion as well as normal use.
A user who wants to remove a watermark can try paraphrasing, translating, changing sentence structure or passing the text through another model. Even without an adversary, routine copy editing may alter the statistical pattern. A document may be assembled from several sources, producing a mixture that is difficult to classify.
The system may also create a new incentive for false reassurance. If a document lacks a Claude mark, a reader could conclude that no AI was involved. That conclusion would be invalid unless the detector covers every relevant model and the text has remained intact. An absent watermark may mean that Claude was not used. It may also mean that another model was used, the mark was removed or the detector failed.
The reverse problem is more serious in disciplinary settings. A positive result can be treated as proof even when it establishes only that a particular model may have processed some text. The difference between “detected” and “demonstrated” is where many automated systems become dangerous.
Anthropic could reduce that risk by publishing a technical specification, a detector for independent testing and clear usage limits. It could report results across document lengths and common transformations. It could test human-written text that has been proofread, translated and summarized. It could disclose performance by language and text type instead of presenting a single average.
It could also make the watermark’s scope visible in the product interface. Users should know when a response is marked and whether the mark applies to generated text only or to any text returned after processing. Enterprise customers should be able to record the operation performed, not just the final document.
That information would not solve authorship. It would make the evidence more honest.
The larger shift is from content detection to process records
The industry has spent years looking for a detector that can identify whether text “was written by AI.” That framing assumes the document itself contains a complete answer. In practice, modern writing is becoming a chain of actions: a person drafts, a model translates, another person checks facts, a colleague rewrites, and a publishing system formats the result.
A single watermark cannot represent that chain.
The more durable approach may be provenance records that describe the process. Cryptographic credentials, edit histories and application logs can record which tool performed which operation at which time. They do not need to infer authorship from the final text alone. They can show that a user requested a translation or that a model generated a particular passage before a human revised it.
Such systems create their own problems. Logs can expose private information. Employees may not want every edit recorded. Different software vendors may use incompatible standards. Records can be deleted or manipulated. A process history is evidence, not an omniscient account of intention.
But it is closer to the question institutions actually need to answer. Did a model generate the argument? Did it only correct spelling? Was a human responsible for checking the claims? Did the user disclose the assistance required by the relevant policy?
Anthropic’s watermark moves the industry toward that debate, even if the technology itself remains underdescribed. It treats model contact as something that can be made visible. The unresolved issue is whether users and institutions will mistake visibility for understanding.
For low-stakes content, that may be tolerable. For a school grade, a job application, a legal filing or a published investigation, it is not. Before a watermark becomes evidence, Anthropic needs to show what its signal detects, how often it fails and what it cannot say.
Until then, the safest interpretation is narrow: Claude may have been involved in producing or processing the text. That is useful information. It is not a verdict on authorship.