From San Digital, who run live AI courses for engineers and business teams.Live AI training from San Digital. See the courses

declawd

How Declawd’s watermark works

Declawd uses a public KGW-inspired watermark. Its seed, cut-off and checker code are public. Anthropic says Claude uses a version of SynthID-Text; the two methods do not interoperate.

This page uses Declawd v2. Its candidate review, minimum length and selection rule were fixed before one seed was sampled. Inspect the exact source commit. The source contract is bound to the published release.

Token pairs

The scanner looks for runs of ASCII letters, with an apostrophe allowed inside a word. Every other character breaks the run. Put an invisible character inside “pressure” and the scanner reads two tokens instead of one.

For each token, it hashes that token, the one before it and the public seed. The rule classifies about one in four possible pairs as marked. Every distinct pair contributes to the score, whether marked or not.

A repeated pair counts only the first time. The score compares the marked count with the expected one-in-four rate. The v2 engine and public pages both require 200 distinct pairs for a verdict. The pages withhold the score below that minimum.

Marked pair rate
1 in 4 token pairs
Cut-off
2.30
Registered v2 minimum
200 distinct token pairs
Public result minimum
200 distinct token pairs
Technical values
Scanner pattern
[A-Za-z]+(?:'[A-Za-z]+)?
Method version
declawd-v2-r2
Public seed
ffbbfdc089c8c5124e1605f69e90300939037f524a7f70e5429472f07a554a55
Registered inputs (SHA-256)
113264ba19926cd039440c4b13e11b030b0708321e6f447186963f466f91516e

Cut-off 2.30

The marked fixture scores 0.77 against the 2.30 cut-off. It misses the cut-off in this run. The fixture result does not set the threshold, and it is retained alongside the corpus results.

We tried cut-offs in steps of 0.05 on 96 calibration passages by 8 authors, using their full text and prefixes ending at 200 and 201 distinct pairs. We kept the lowest cut-off that left no more than 2% above it in each of those three groups separately.

We then used it on 96 passages by 8 other authors. No author is in both sets. 6 passages crossed: 6 of 96 passages, or 6.25%, which is above the 2% we registered before the run. We left the cut-off unchanged after evaluation. V2 reuses the published v1 corpus and author split. The experiment uses a new seed against known material, and both halves were public before this registration.

The full calibration set returned: 1 of 96, or 1.04%, with a highest score of 3.00. Across both sets, 7 of 192 older human passages crossed the cut-off, and the highest score any of them reached is 3.00.

Calibration passages hold 200 to 622 distinct pairs, and the evaluation passages hold 200 to 638. V2 also measures the 200 and 201-pair prefixes to test the registered minimum directly. Prefixes ending at 199 pairs return insufficient text. Issue #17 records the original grammar defects and the v1 minimum of 120 pairs. V1 remains archived with its original inputs, seed and results.

The table keeps each length group separate. Prefixes from one passage overlap, so they are repeated views of the same material. They are not additional independent passages.

Scroll horizontally to see all columns

V2 results at the registered length boundary
TextCalibration crossingsEvaluation crossingsEvaluation intervalUnavailable prefixes
199 distinct pairsInsufficient text (96 passages)Insufficient text (96 passages)No verdict0 of 96 calibration / 0 of 96 evaluation
200 distinct pairs1 of 96 (1.04%)1 of 96 (1.04%)0.18% to 5.67%0 of 96 calibration / 0 of 96 evaluation
201 distinct pairs1 of 95 (1.05%)1 of 94 (1.06%)0.19% to 5.78%1 of 96 calibration / 2 of 96 evaluation
Full passages1 of 96 (1.04%)6 of 96 (6.25%)2.9% to 12.97%0 of 96 calibration / 0 of 96 evaluation

A Wilson calculation puts the 95% interval at 2.9% to 12.97%, but it assumes that the passage results are independent. That may not be true here because several passages come from each author.

Limits

  • The scanner recognises A–Z, a–z and an apostrophe inside a word. That is the entire alphabet. Text with no A–Z words is refused outright, which is the safe case. Everything else is harder. Latin text with accents is split at every accent, so French, German, Spanish and similar writing reaches a verdict built from fragments. Mixed text is scored on its Latin fragments alone, so a page written mostly in another script can still return a verdict from its names, quotations or reference list. The cut-off was measured on English prose. Do not read a result for any other language.
  • The home page uses prepared passages. The checker scores pasted text in your browser and sends no text over the network. Neither page accepts files. The v2 registration sets the 200-pair result minimum before seed sampling.
  • The repair reverses only the zero-width space and Cyrillic lookalikes used on this site.
  • Do not carry this score or the 2.30 cut-off over to Claude or another watermark.
  • The worked SynthID-Text fixture uses a separate fixed token trace and has no detector cut-off.

Give the room five minutes with the four steps. Then put a real piece of work on screen and ask where it came from, what the model did, what the person changed and what happened next. Finish with the awkward bit: which of those facts could you prove today, and which are worth recording next time?