Kirchenbauer and others
Introduces the green-list test used here and recommends counting repeated n-grams once, so repetition cannot drive the score.
A Watermark for Large Language ModelsVendor statements, legal text, research papers, books, dated reports and independent tests: the main sources behind the argument on Declawd are on this page. Not the operational references, though. Those live where they are used, so the hosting and data-protection material sits on Privacy and the company registration in the footer. Baseline checked 12 August 2026.
Baseline source review: 12 August 2026. Later checks are dated by section.
Anthropic says finding a supported Claude mark can mean that Claude processed the content, not that Claude wrote it. Proofreading, translation and format conversion can all produce a Claude mark, though the carrier and detectable density depend on what Claude changed.
Not finding one is inconclusive too. Claude or another model may still have been involved.
Anthropic now identifies its text method as a version of SynthID-Text. It says a key and preceding words alter selection among suitable next words, with no hidden characters added. Its Help Centre lists four models with text watermarks, Fable 5.1, Mythos 5.1, Opus 5.5 and Opus 5, says every older model will be covered by 2 December 2026, and offers text detection in private preview. Its public file checker reads C2PA credentials. Declawd has not accessed the private detector.
Checked 24 September 2026. The 1 September Platform release notes identify Fable 5.1 and Mythos 5.1 as marked and describe Content Credentials on supported generated files. The 22 September Opus 5.5 entry does not mention marking, so Opus 5.5 and Opus 5 appear as marked only in the Help Centre table. The public file checker does not test text watermarks.
Anthropic, How Claude’s text watermark works. Anthropic, How Claude marks AI-generated content. Models API documentation. Platform release notes. Claude Content Checker
Anthropic is not the only company bound by the Code of Practice. Reading its rivals side by side shows how uneven the picture still is, and how little any of it helps a person judging one piece of text.
OpenAI marks supported images with C2PA metadata and a SynthID watermark, and supported audio with a SynthID watermark. It runs a public verification page and an API for the provenance it adds. For text it says only that its goal is to widen that provenance to every kind of output as standards mature. Anthropic now offers its text-watermark API in private preview, alongside a public C2PA file checker.
One detail is worth pausing on. OpenAI uses SynthID, which Google built, and Anthropic names a version of the same family for Claude text. Three of the larger providers now lean on one lineage, which is a reason to explain how it works and a reason none of their detectors reads another’s output.
Google says the Gemini app and web experience have marked text with SynthID-Text since 2024. Its public verification route is to upload an image, video or audio clip to Gemini and ask. Text is not in those instructions, and its detector portal is limited to vetted journalists and researchers. The open developer implementation watermarks and detects with a configuration and key the operator controls, so it verifies an operator’s own marks and not Google’s production ones. What Google’s official pages do not document is the production API: neither whether and since when Gemini API text output carries the mark, nor any public way to verify production text. On Google’s developer forum a participant whose profile title reads Google answered on 5 August 2026 that API text was not watermarked, then corrected that on 19 August to say it is, including text from AI Studio. The forum marks the account as neither staff nor moderator, and a follow-up asking since when has had no answer. A public cleaner’s documentation says the API text watermark was retired in August 2026, which conflicts with that correction. Declawd records the API’s marking history as undocumented rather than settled. A change Google shipped in August 2026 turns the visible badge in the Gemini app on or off; it does not touch the hidden watermark, and the two are easy to confuse.
None of this lets Declawd test Claude. Each verifier reads its own maker’s output alone. OpenAI’s page says as much, and finding no mark is not a finding that none is present. OpenAI sources rechecked 18 August 2026: the developer guide was read live, while the two help-centre pages block automated reads and were confirmed through search results. Google sources checked 24 September 2026.
Article 50(2) says providers must mark AI-generated or manipulated output in a machine-readable form and make it detectable as artificially generated or manipulated. It says the technical solution must be effective, interoperable, robust and reliable as far as technically feasible.
The Act does not currently prescribe one method, though a 2026 amendment lets the Commission set common rules if a code of practice proves inadequate. The duty does not apply to the extent that a system assists with standard editing, does not substantially alter the deployer’s input or its meaning, or is used under legal authority to detect, prevent, investigate or prosecute crime.
Anthropic says models launched in the EU from 2 August 2026 support marking at launch, with the marks applied wherever Claude is offered. Providers of systems that generate synthetic audio, images, video or text and were placed on the market before that date have until 2 December 2026 to comply with Article 50(2), rather than with the Act as a whole.
Consolidated Regulation (EU) 2024/1689, as amended in 2026. The consolidated text is a documentation tool and has no legal effect; the Official Journal version governs.
Introduces the green-list test used here and recommends counting repeated n-grams once, so repetition cannot drive the score.
A Watermark for Large Language ModelsThe paper describes the sampling method and reports a live experiment covering nearly 20 million responses. Anthropic says Claude uses a version of the method, but the paper does not describe Claude’s production key, configuration or detector.
Scalable watermarking for identifying large language model outputsTests character edits against spell correction, OCR, Unicode normalisation and deletion of unusual characters. Adaptive attacks still worked after each defence.
Character-Level Perturbations Disrupt LLM WatermarksA versioned preprint. It reports watermark removal under stated conditions, plus detector errors, in the KGW, Unigram and MarkLLM SynthID-Text configurations it tested. Declawd has not reproduced the experiments.
Checked 13 August 2026.
AI Watermark Evidence Fails Forensic Readiness, preprint v1The authors report learning approximate rules for tested watermark schemes and using them for removal and spoofing. Declawd has not reproduced the experiments.
Checked 13 August 2026.
Watermark Stealing in Large Language ModelsZhang and others report an impossibility result under the quality-oracle and perturbation-oracle assumptions the paper sets out. Declawd has not reproduced the analysis.
Checked 13 August 2026.
Watermarks in the SandA theoretical and empirical analysis reports that SynthID-Text’s mean detection score is inherently vulnerable to added tournament layers, and demonstrates a layer-inflation attack that breaks it while a Bayesian score holds up better. This concerns the published scoring mechanism, not Anthropic’s production configuration; Declawd has not reproduced it.
On Google’s SynthID-Text LLM Watermarking SystemA COLM 2026 paper reports that existing token-level watermarking is inadequate on constrained output such as translation, summarisation and code, where there are few word choices to carry a mark, and proposes a sequence-level method it says recovers up to 28 per cent in detection F1. It tests open schemes, not Claude.
Semantic Differentiation for Watermarking Low-Entropy Constrained GenerationA preprint audits six open watermarking schemes on three open-weight generators across eleven languages, four scripts and eight typological families, and reports detection and quality gaps that fall mainly between language families rather than within them. It tests open schemes, not Claude’s production one, and it is a reason not to read any vendor’s quality or robustness claim as holding across languages. Declawd has not reproduced it.
Checked 21 August 2026.
Auditing Cross-Lingual Fairness in Language Model Watermarking, preprint v1A preprint shows that the share of a text that is edited is an incomplete robustness measure: at one and the same retention rate, the surviving detector statistic can be half the original, a quarter of it, or nothing, according to where the edits fall. Its experiments use schemes the author re-implemented with his own keys and a 0.5B open model, and it measures no production scheme. Declawd has not reproduced it.
Checked 21 August 2026.
Linguistic Holonomy and Statistical Watermarks, preprint v1An amendment to an open SynthID-Text evaluation reports a same-model Bayesian detector at 73.58 per cent sensitivity and 2.78 per cent false positives on texts of 200 tokens or more, and works through what that means at an illustrative 1 per cent prevalence: a positive predictive value of 21.1 per cent. The 1 per cent figure is an illustration, not a deployment estimate. The detector, key and Qwen model are the authors’ own, not Google’s production system. Declawd has not reproduced it.
Checked 8 September 2026.
Detector score family amendment, xlr8harder/synthid at revision b3adb0funsynth-id reports bit-for-bit conformance against the public SynthID-Text reference, key recovery for all nine reference layers in 1.8 seconds, and rewrite results under its documented conditions. It calls itself a research testbed, not a product. Dependencies, model files and every result input are not all frozen in this capture. Declawd has not rerun it, and it does not test a production vendor system.
Checked 24 September 2026.
unsynth-id at revision 9cfef6eTrellisMark publishes an experimental watermark that encodes a 30-bit user address and uses exact Viterbi decoding across the full address space. The repository warns that it demonstrates a surveillance capability and says no lab has announced deployment. It is an independent Qwen research system, not evidence about Claude.
Checked 24 September 2026.
TrellisMark at revision 3f0157dSources first checked 13 and 14 August 2026. John Wang and the two press reports were rechecked 24 September 2026. Later additions are dated on their cards.
John Wang reports finding no anomalous hidden Unicode or whitespace encoding in a 7.2-million-character corpus. The post does not document a class-for-class comparison with this site’s registry or publish the underlying response corpus.
How Claude watermarking probably worksThe same report describes negative green/red-list constrained-choice permutation tests. Its separate SynthID-style probes found prompt effects without identifying a watermark, which is not a finding that the tested models are unmarked.
Read the test detailsA later study builds on the same Red-Green test and shows the test number as spaced digits. It reports a keyed pattern in Fable 5.1 output tested on 2 September, with a corrected p-value of 0.006 for a 15-digit number and close to zero for a 40-digit one, and reports Sonnet 5 as null. It tests whether a model’s sampling carries a keyed bias across many answers, so it cannot say whether one passage is marked. Declawd has not rerun it.
Checked 24 September 2026.
KarenSpinner/watermark-detection-study at revision 8838e0dArs Technica reported that Anthropic’s 13 August reply did not answer questions about detector timing, error-rate testing or editing exemptions. Anthropic now documents a private preview of its text-watermark API. Declawd has not accessed that detector or measured its errors.
Ars Technica reportTechCrunch, on 15 August, reported Anthropic’s further detail that a detection API is planned with no date, a complete rewrite removes the mark, and code carries so little that only comments are likely to. It still records no error rate, threshold or minimum length. Its planned-API wording is superseded by Anthropic’s September private preview. Checked 24 September 2026.
TechCrunch reportAn independent repository replicates SynthID-Text with the official Hugging Face implementation and its own key on Qwen2.5-7B, then attacks it with a second open model. It reports detection falling from 100 per cent on unedited English text to 5 per cent after paraphrase, 16 per cent after round-trip translation through Chinese and 4 per cent after a self-information rewrite. Its README says none of this is measured against Claude’s real watermark, and the authors did not test Anthropic’s private detector. The repository provides a way to run tests. Only dated runs such as these count as results, and Declawd has not rerun them.
Checked 8 September 2026.
natzir/synthid-text-watermark-attacks at revision 57ae015Three bounded GitHub searches for public repositories created from 15 August 2026 to 24 September 2026 returned 1163 distinct matches. The capture holds 1137 READMEs and records 26 gaps. Automated screening covered every captured README. The review records which documents and source files were read closely.
Some repositories claim Anthropic API integration or allow an operator to configure a detector URL. The inspected material did not demonstrate a working provider-backed integration or removal verified against one. Declawd has not accessed Anthropic’s private text detector.
Captured 24 September 2026. The next sweep starts again at the fixed 15 August discovery floor.
API captureIssue #17 reproduces two candidate slots that do not fit their sentence frames and a 120-pair result floor below the 200-pair minimum in the calibration and evaluation evidence. The public site now withholds scores and verdicts below 200 pairs. The registered v1 inputs remain unchanged. V2 corrects the candidate frames, registers a 200-pair minimum and checks both sides of that boundary with a fresh seed. The method reports its results separately from v1.
Reviewed 24 September 2026.
Declawd issue #17These studies cover classifier and perplexity detectors, not watermarks. They are here because people can turn either kind of score into a judgement about a writer.
In 2023, seven widely used detectors returned a mean false-positive rate of 61 per cent on essays by non-native English writers, against near-perfect accuracy on a native-speaker comparison set. A 2026 study found no systematic bias across three detector families in Czech, but the detector it tested still showed a 23.1 percentage-point gap in false-positive rate between the non-native and native English sets used in 2023. The result depends on the detector, language and population.
A C2PA manifest can carry assertions about an asset’s source and editing history inside a signed claim. A validator checks the signature, the credential that signed, and how the manifest is bound to the asset, among other things.
The signature does not make every assertion true, and it cannot say who did the work. The credential can identify whoever signed, who may not be the maker of the asset at all. Attributing work to a person requires its own identity assertion, unless a tightly run workflow stands in for one.
Declawd v0.1.1 includes signed PNG and JPEG fixtures and an independent removal oracle, and the C2PA walkthrough follows one test image through four validator states, with every asset and the unedited validator output published alongside it. Anthropic now offers a public C2PA file checker and text-watermark detection in private preview. The file checker explicitly excludes text detection. Declawd has not accessed the private API.
One distinction worth knowing: the signer vouches only for assertions the claim generator created itself. Assertions gathered from elsewhere get no such vouching, yet the signature covers both kinds. A claim generator, by the way, means hardware or software, never a person. Two kinds of binding exist as well. A hard binding hashes some or all of the asset’s bytes. A soft binding is looser; it can reconnect an altered asset to a manifest held somewhere else.
Project Gutenberg lists the 16 books below as public domain in the United States. Eight authors supplied the 96 passages used to set the cut-off; the other eight supplied the 96 passages used to test it. No author is in both sets.
Scroll horizontally to see all columns
| Author | Work | Source |
|---|---|---|
| Adam Smith | The Wealth of Nations | Project Gutenberg 3300 |
| Benjamin Franklin | Autobiography | Project Gutenberg 20203 |
| Charles Darwin | The Origin of Species | Project Gutenberg 2009 |
| David Hume | An Enquiry Concerning Human Understanding | Project Gutenberg 9662 |
| David Ricardo | Principles of Political Economy | Project Gutenberg 33310 |
| Henry David Thoreau | Walden | Project Gutenberg 205 |
| John Locke | An Essay Concerning Human Understanding | Project Gutenberg 10615 |
| John Ruskin | The Crown of Wild Olive | Project Gutenberg 26716 |
| John Stuart Mill | On Liberty | Project Gutenberg 34901 |
| John Tyndall | Fragments of Science | Project Gutenberg 24527 |
| Michael Faraday | The Chemical History of a Candle | Project Gutenberg 14474 |
| Ralph Waldo Emerson | Essays | Project Gutenberg 16643 |
| T. R. Malthus | An Essay on the Principle of Population | Project Gutenberg 4239 |
| Thomas Henry Huxley | Evidence as to Man's Place in Nature | Project Gutenberg 2931 |
| Thomas Paine | Common Sense | Project Gutenberg 147 |
| Thorstein Veblen | The Theory of the Leisure Class | Project Gutenberg 833 |
Public domain in the United States. Read the Project Gutenberg licence and terms.