What public cleaners change
Published ; updated
The Declawd measurements in this article use archived v1. Their original figures and evidence are preserved. The current method and interactive pages use v2, so links from this article to them show v2 figures.
Anthropic describes Claude's text mark as a version of SynthID-Text and says a complete rewrite removes it. Its text detection API is now in private preview for eligible organisations. Declawd has not accessed it, and the independent tools reviewed here have not demonstrated removal against it. Software can establish narrower results: selected Unicode characters are gone, an embedded C2PA store is absent, or the wording changed. Testing the production text mark requires compatible detector access.
Names
The term "watermark" covers several distinct technical carriers. A tool that handles one of them may have no effect on another.
Scroll horizontally to see all columns
| Carrier | What can be changed | What a successful operation establishes |
|---|---|---|
| Selected Unicode characters | Remove or replace exact code points | Those selected characters are absent from the output |
| Embedded C2PA manifest | Remove the manifest store from a supported file | That supported embedded store is absent from the output |
| Statistical token-choice pattern | Reword some or much of the passage | Word choices changed. Anthropic says a complete rewrite removes its mark, but Declawd has not accessed its detector in private preview |
| Pixel or soft-binding channel | Re-encode or alter image content | Nothing conclusive unless the matching resolver or detector is tested |
Unusual Unicode is not evidence of AI generation. Joiners shape Persian, Arabic, Urdu, Indic text and emoji. Direction controls make bidirectional text readable. Variation selectors can determine whether a symbol appears as text or emoji. Spaces and combining marks can also be meaningful. Removing them by category can damage ordinary human writing while leaving a statistical watermark untouched.
C2PA metadata works differently. A standard C2PA manifest contains signed assertions and a hard binding to an asset. The C2PA specification expressly allows its metadata to be removed. The specification also defines soft bindings that can reconnect an asset to a remotely held manifest after embedded metadata has gone. Removing an embedded manifest does not prove that every provenance channel has disappeared.
Current cleaners
Reproduced observation, captured 24 September 2026 at 15:22 UTC. Three GitHub searches covered public repositories created from 15 August to the capture time. All three explicitly searched names, descriptions and READMEs. The Claude, Anthropic and broader text-watermark queries returned 43, 9 and 1,144 matches, with 1,163 distinct repositories in their union. The broader query now passes the 1,000 results a single search will return, so it was split by creation date into two slices that together account for every match. The capture contains 1,156 default-branch heads and 1,137 READMEs pinned to those heads. It records 26 README capture gaps, including seven empty repositories. The capture records every query, slice and repository.
Names and descriptions were reviewed across the set. Automated screening covered all captured READMEs. Where a README had not changed since 8 September, its judgement from that review was kept. The 437 new repositories with a captured README, the 50 whose README changed and the 2 whose earlier README gap has since been captured were screened again, with selected documentation read closely. The review report distinguishes those levels of review and records duplicates, capture gaps and which judgements were carried forward. The inspected material did not demonstrate a working Anthropic detector integration or removal verified against one. This is a finding about the material examined at those revisions. Declawd has not accessed Anthropic's private detector.
Some repositories do claim provider integration. GhostMark at revision bacaa88 offers an optional Anthropic API route. Its source sets a default endpoint and response fields that this review could not verify against an official contract, while allowing the endpoint to be changed. The reviewed material supplies no verified official contract or successful provider result for those assumptions. anthropies and markprobe also accept operator-configured detector URLs. The presence of HTTP transport establishes that code can send a request. Compatible detection still needs its own evidence.
A Claude watermark benchmark, FelixDes/claude-watermark-benchmark, read at revision 1215052 on 8 September, reported probes on Fable 5.1, described its results as null or confounded, and explicitly awaited official detector ground truth. By 24 September the repository and that revision returned HTTP 404, and no archived copy was found, so the reading is recorded as it stood on 8 September and can no longer be checked against the source. A later black-box study reports a keyed pattern in Fable 5.1 output. Neither result has been reproduced by Declawd.
The earlier sweeps matched 176 repositories on 18 August, 292 on 22 August and 434 on 27 August. Their dated records remain in the ledger and the 27 August capture and review. The September review uses a narrower conclusion than those earlier absence claims, because it records unverified integration claims as such.
Public implementations remain useful for reproducible research. An independent detector implements six published schemes and requires an operator-supplied key. The unsynth-id study below and the TrellisMark attribution research concern open research configurations. Their results do not measure Claude's production mark.
demark at the reviewed revision has its output-safety discipline in order. Inspection never writes. Ordinary cleaning produces a new file and leaves the original alone. There is an in-place mode, but the user has to ask for it, and it takes a backup unless that is switched off too. The tool re-inspects whatever it produces, and it keeps verified structural cleaning well apart from unsupported statistical and pixel-domain claims.
Its default cleaning scope is broader than Declawd's. According to its documentation, an ordinary clean removes joiners, bidirectional controls, variation selectors and unusual spaces, and removes broad image metadata families rather than only an embedded C2PA store. Those structures can carry language, presentation, attribution or accessibility meaning. Its PolyForm Noncommercial licence is also a use constraint: we can study and cite the published design, but Declawd does not copy, depend on or execute it for a commercially published benchmark without permission or legal clearance.
watermarks-remover at revision 43959fac documents v0.6.0. It adds plugin, hook, service, multi-format and benchmark infrastructure. Those are substantial product changes, but the evidential boundary is unchanged. At that captured revision, its Claude detector is a placeholder awaiting Anthropic's API. Its MarkLLM checks remain same-configuration-only and are not a vendor oracle. The repository contains a benchmark harness but no committed result that proves a vendor mark was removed. It still says Google retired API text watermarking in August 2026. The two Google official pages reviewed here do not document the production API's history, and a participant on Google's developer forum corrected an earlier answer on 19 August to say that API text is watermarked. The forum JSON still contained five posts on 8 September, with the account marked as neither staff, admin nor moderator and the follow-up asking since when unanswered. Declawd therefore records the API's marking history as undocumented rather than settled. Without compatible detector access and secret material, no result from this tool certifies what a vendor detector would return.
The separately maintained Leutenegger/watermarks-remover repository reviewed on 22 August returned not found on 27 August. Its pinned revision remains part of the dated historical record, but it is no longer presented as a current tool.
Studies
Reported result. Tamim and Khan report 846 valid paraphrase runs in a preprint pinned at v1. Among texts where the mark was initially detected, one meaning-preserving paraphrase removed it in 100 of 100 KGW runs, 41 of 41 Unigram runs and 58 of 59 runs, or 98.3 per cent, for the MarkLLM implementation of SynthID-Text. Before any attack, false-negative rates were 70 to 83 per cent. The tested SynthID configuration flagged 5.4 per cent of paraphrased human-written controls. The authors assess the results against Daubert, a US federal evidentiary framework, and NIST SP 800-86, a forensic process, and conclude that detection in the tested configurations is not ready for evidentiary use.
Reported result. Jovanovic, Staab and Vechev report learning approximate rules for each tested watermark scheme at a one-time query cost under 50 dollars per scheme, using January 2024 ChatGPT API prices. The learned approximation supported removal and spoofing, creating false-attribution risk.
Reported result. Zhang and others report that under the paper's quality-oracle and perturbation-oracle assumptions, strong watermarking is impossible.
Reported result. A theoretical and empirical analysis of SynthID-Text reports that its mean detection score is inherently vulnerable to added tournament layers, and demonstrates a layer-inflation attack that breaks that score while a Bayesian score holds up better. This is about the published scoring mechanism, not Anthropic's production configuration, and Declawd has not reproduced it.
Reported result. A preprint on edit placement shows that the share of a text that is edited is an incomplete robustness measure: at one and the same retention rate, the surviving detector statistic can be half the original, a quarter of it, or nothing, depending on where the edits fall. Its experiments use schemes the author re-implemented with his own keys and a 0.5B open model, and the paper says it measures no production scheme. Declawd has not reproduced it.
Reported result. An independent reproduction at revision 57ae015 replicates SynthID-Text with the official Hugging Face implementation and its own key on Qwen2.5-7B, then attacks it with a second open model. It reports detection falling from 100 per cent on unedited English text to 5 per cent after paraphrase, 16 per cent after round-trip translation through Chinese and 4 per cent after a self-information rewrite. Its README says none of this is measured against Claude's production watermark. Declawd has not rerun it.
Reported result. unsynth-id at revision 9cfef6e reports bit-for-bit conformance against the public SynthID-Text reference. Its best meaning-preserving rewrite covers 8 documents and 3 seeds, moving mean-g z from 16.38 to 1.99 while preserving every numeral and named entity in 87.5 per cent of documents. The README also warns that mean-g near 2 is suppressed rather than null and may remain detectable on a long document. It reports recovering all nine reference layers in 1.8 seconds. The repository calls itself a research testbed, not a product. Its dependencies, model files and every result input are not all frozen in this capture, and Declawd has not rerun it. None of the result is about Claude or Google's production configuration.
Unknown. Results against published schemes, including a MarkLLM implementation of SynthID-Text, do not establish the behaviour of Anthropic's version, key or detector.
Feedback
Official statement, checked 24 September 2026. Anthropic's technical post, updated on 1 September, now describes text detection in private preview for eligible organisations. The Help Centre now lists four models with text watermarks, Fable 5.1, Mythos 5.1, Opus 5.5 and Opus 5, and says every model released before 2 August will be covered by 2 December. Anthropic also provides a public file checker for Claude-issued C2PA credentials, which explicitly does not check text. Declawd has not accessed the private text API. The available access routes and remaining unknowns are covered in the announcement article.
Inference, not evidence. A broadly accessible detector returning useful feedback could support adaptive retries against itself. Anthropic says it plans to expand private-preview access over time, without publishing a general-access date. The announcement article explains the Code's permitted access routes.
Design
Declawd follows an inspect-first workflow, writes to a separate file and checks the output. Its report names every untested channel, and its cleaning scope is narrower.
Inspecting text lists every code point along with its position. Nothing is selected by default. The user chooses an exact character or a narrowly named class, sees the proposed action and keeps the original. The public Declawd v0.1.1 CLI removes only a supported embedded C2PA store from PNG or JPEG while preserving image data and unrelated metadata. The CLI does not offer blanket Unicode normalisation, automatic homoglyph folding, broad metadata stripping or an in-place mode.
Declawd reports only what it checked. It can state that the selected structures were removed and that no supported embedded carrier remains. Statistical, soft-binding and pixel-level marks were not tested, and the report lists them as untested. It never says "Claude watermark removed".
Misuse
A stray hidden character can break search. It can break source control, or some downstream system that chokes on the byte. Embedded provenance brings its own problem, since it can expose a workflow or a tool choice its author never meant to publish. In short, people often hold perfectly sound reasons for inspecting or altering a file of their own. Running the cleaner is a direct way to see those limits.
Where does trouble start? With a cleaner that promises more than it can measure. Picture a one-click service that rewrites arbitrary prose and stamps the result "human": at that point an explanation has become an evasion claim. Declawd publishes something narrower, a controlled attack on its own disclosed watermark alongside a user-directed word editor and explicit structural cleaning. No synonym suggestions, and no calls out to another model. It claims to defeat no vendor detector either.
Use is limited to content you own or are authorised to modify. Finding or removing a listed character does not show that AI was involved, and a cleaned file establishes nothing about authorship.
Gaps
Anthropic now identifies marked models and offers text detection in private preview. Compatible access, a documented decision rule and error-rate evidence are still needed for Declawd to run and interpret a production before-and-after test. The documentation and source inspected on 24 September did not demonstrate removal verified by Anthropic's detector.
Try the educational cleaner, or read the controlled experiment.