From San Digital, who run live AI courses for engineers and business teams.Live AI training from San Digital. See the courses

declawd

Six word changes against our published watermark

Published ; updated

The Declawd measurements in this article use archived v1. Their original figures and evidence are preserved. The current method and interactive pages use v2, so links from this article to them show v2 figures.

Declawd publishes a small statistical watermark with its seed and detector. With those details public, we can change selected words, score the whole passage after each change and record exactly when it crosses the published cut-off. Anthropic now says Claude uses a version of SynthID-Text, but has not published the production key, compatible detector or decision rule needed to score it. This experiment attacks Declawd's educational method, not Claude's watermark.

Scroll horizontally to see all columns

The two watermark methods are not interoperable
PropertyDeclawdClaude, according to Anthropic
MethodPublic KGW-inspired pair scorerA version of SynthID-Text
Key and detectorPublishedProduction key and detector unavailable
Result on this pageReproduced against a fixed cut-offCannot be calculated or inferred

Start

Profile declawd-v1 works on ASCII word tokens. Each distinct pair of the previous and current token gets scored, with a repeated pair counting once only. A published seed fixes which pairs are green. The expected green fraction is one quarter. The frozen v1 contract returns no engine verdict below 120 distinct scored pairs. After issue #17 identified that every passage used to choose or test the cut-off has at least 200 pairs, the public site began withholding scores and verdicts below 200. The registered v1 contract remains unchanged so the published evidence stays reproducible.

The issue also tests the short range directly. It truncates all 192 human corpus passages to the shortest prefix that reaches 120 pairs. Four score above the 1.80 cut-off. The highest, a Benjamin Franklin passage at 3.16, is above the marked fixture at 2.99. The site's 200-pair floor keeps public verdicts inside the measured range. It does not repair or reinterpret the frozen engine.

The prepared flow-meter passage contains 398 raw tokens and 358 distinct scored pairs. Of those pairs, 114 are green. Its score is 2.9904, above the cut-off. The profile, passage template, authored candidate slots and scoring rules were frozen before this removal experiment was designed.

Two authored v1 candidate slots also fail in their sentence frames. The marked fixture selects wording that produces "due the delay" and "carry with a run". Those awkward phrases remain visible because changing the registered template in place would break the evidence binding. A future v2 needs to render and validate every candidate in context before registration, then use a fresh seed.

The exercise uses a symmetric keyed hash with a published seed, not asymmetric cryptography. The published value is the seed.

Substitutions

Each replacement below comes from the same authored candidate slot as the word it replaces. We selected changes with access to the detector, then reviewed the result as an ordinary piece of technical writing. The table reports scalar offsets in the original passage, before any replacement changes later positions.

Scroll horizontally to see all columns

Detector-aware substitutions applied to the prepared passage
StepOriginal scalar offsetChangeGreen pairs after changeScore after change
StartNoneNo change1142.9904
175sample to measure1122.7463
2175first value to figure1102.5021
3295holding to resting1082.2580
4371first passes to repeats1062.0139
5631falls to comes1041.7698
6994lies to sits1021.5257

Watch the repeated words. There is a second "value" at original offset 327 and a second "passes" at 1030, and neither gets touched. Offsets drift too. The first five changes add three characters between them, so by the time the last replacement is applied it begins at offset 997. That is why the table records original offsets: every target stays unambiguous.

After five changes the score stands at 1.7698, under the 1.80 cut-off but only by 0.0302. That margin is less than one scoring step, because with 358 effective pairs a single green-pair change moves the z score by about 0.1221 here. The sixth change takes the score to 1.5257, which leaves 0.2743 below the cut-off.

Pairs

Every scored pair is built from the previous token and the current one. So a single replaced word can shift two pairs at once, the pair ending at that word and the pair starting from it. One visible word changes. The green count drops by two. That is exactly what each selected change above does.

Recomputing the entire passage is not optional for a guided editor. A replacement can change the contexts either side of it, and a repeated pair only counts at its first occurrence, so nothing local can be trusted on its own. Treat a highlight as a place to investigate. It is no promise that the score will fall by any fixed amount.

In the cleaner, the replacement comes from the user. The tool previews the new whole-passage score and applies nothing until the user confirms it. It generates no synonyms and rewrites no text silently. And once an edit lands, the context table gets rebuilt, since the previous ranking cannot be assumed to still hold.

Result

Reproduced observation. Six human-reviewed, detector-aware substitutions move this fixed passage from 2.9904 to 1.5257, with the effective pair count holding at 358 throughout. The registered candidate slots constrain what wording is available. Take the published profile and the reference implementation and you can reproduce the exact result yourself.

The experiment does not establish that meaning is mathematically unchanged. Semantic review is a human judgement recorded for this fixture. Nor does the experiment show that six edits will clear another passage, another seed, another detector or Claude's mark. A different word can increase the score, change the effective pair count or damage the writing.

Crossing the cut-off does not make the result "human". In the frozen evaluation, 2 of 96 passages by named public-domain authors crossed this site's threshold. The score answers one narrow question about one published pattern. It does not answer who wrote a passage or how they worked.

Encoding

Inference, not evidence. A public post on 17 August, still carrying no measurements or replies when its API record was checked on 8 September, suggested that asking a model to answer in base64 and decoding the result afterwards would shake off a distribution-based watermark. Generation carried out in the encoded form is something a fixed-template watermark like this one cannot be made to do, so Declawd cannot set that case up and neither confirms nor refutes it. What it can do is measure the transform against its own published method.

Reproduced observation. Base64 used as transport changes nothing. Encode the marked passage and decode it again and the bytes come back identical, scoring 2.9904. Score the base64 text itself, without decoding, and it reads as 260 tokens with 255 distinct pairs, 53 of them green, for a score of -1.5547, below the 1.80 cut-off. Re-encoding rewrites the words the scorer reads, so it counts different pairs, not the same pairs made to disappear. The bundle behind these numbers, carrying the scorer, the profile and the marked passage, is published to rerun: the reproduction script, the marked input, the hash manifest and the dated report.

Modes

In the controlled fixture every choice is inspectable. Reproducibility comes from what was pinned down in advance: the registered passage and its existing candidate slots, plus a fixed expected result.

The guided mode lets a person explore their own text, provided it is long enough, and it comes with fewer guarantees. Below 200 distinct pairs the public site shows no score, verdict or attack table. At or above that presentation floor the editor can show contributing contexts and preview a replacement the user has written. What it cannot do is promise that an acceptable alternative exists, or decide on the user's behalf that a semantic change is harmless.

Publishing a detector makes detector-aware research possible. An automatic rewriting service would be a different product with a much broader claim.

Try the controlled experiment and guided editor, or read what current cleaners can and cannot change.