No score proves authorship
A watermark test shows only that text matched a set of rules. It says nothing about the drafting and review behind the finished work.
Anyone who wants to remove one can rewrite the text, pass it through another model or change it in several other ways. Human writing can cross the same cut-off by chance.
If you have to decide later, you will need the work history. The finished file will not be enough.
The watermark here is ours and published in full. Anthropic says Claude uses a version of SynthID-Text; Declawd does not reproduce or interoperate with it. This page starts with prepared text; the checker accepts pasted text in your browser. Do not use any result to judge someone.
What above and below threshold mean
Above threshold means the score is greater than 2.30; below threshold means it is not. That is all the result says. It cannot tell you who wrote the text or whether a model had any part in it.
How is the score calculated?
A higher number is a closer match to this watermark, not stronger evidence about authorship. The calculation uses distinct token pairs, and repeated pairs count once.
If a passage contains fewer than 200 distinct pairs, the public pages show “Not enough text” and withhold the score and verdict.
A walkthrough of the published watermark
We add it to prepared text, change a few characters, rewrite the passage and then run the same test on older human writing.
Technical details
- Threshold
- 2.30
- Expected marked share
- 1 in 4 token pairs
- Profile
- declawd-v2-r2
The instructions before and after marking
Here are the same instructions twice. 12 of the 430 words change in the second version and the score moves from -0.62 to 0.77. Every candidate in this v2 template was reviewed in its sentence before the seed was drawn. Marking selects among those candidates, and the resulting passage is below threshold. The marked fixture misses the registered cut-off in this run. We keep that outcome instead of choosing another seed.
Each score uses the full prepared passage. The cards show their opening lines. Read the method.
- Full-passage result
- Below threshold
- Full-passage score
- -0.62
Set the reference flow to 40 litres per minute before you record the first reading on the raw sheet. Hold the line at that rate for two minutes so the pressure can settle.
- Full-passage result
- Below threshold
- Full-passage score
- 0.77
Set the reference flow to 40 litres per minute before you record the first value on the raw sheet. Hold the line at that flow for two minutes so the pressure can settle.
Small edits, then a repair attempt
Change one character and watch the score. One option hides a character inside a word; another swaps in a Cyrillic letter that looks Latin. We can put those two changes back in this prepared English passage. Delete a letter, though, and there is no safe way to infer what is missing.
The last option targets every tenth word. It changes 39 of 42 targets, an editing rate of 0.09, and moves the score from 0.77 to -0.23, below the 2.30 cut-off. It chooses words without knowing where the pattern sits, so this result tells us only what this one unguided attempt did.
Apply one change to the marked passage
Nothing has been changed. Choose an option above.
Set the reference flow to 40 litres per minute before you record the first value on the raw sheet. Hold the line at that flow for two minutes so the pressure can settle.
Technical detail
No change applied. This is the watermarked passage.
Choose a change to see whether this page can repair it.
378 distinct token pairs. Current score: 0.77. Nothing has been changed or repaired.
Added an invisible character inside ‘pressure’. The scanner now reads it as two words.
Hold the line at that flow for two minutes so the preszero-width spacesure can settle.
Can the original text be restored?
Yes. Remove the hidden character and the original passage returns exactly.
Technical detail
Inserted Zero Width Space, U+200B inside the word “pressure” at position 153. Nothing visible changed. The word now scans as two tokens.
The normaliser reversed the allow-listed character change using the submitted text alone.
Scroll horizontally to see all columns
| Position | Submitted value | Normalised value | Reason |
|---|---|---|---|
| 153 | Zero Width Space, U+200B | Removed | Allowed character for this passage |
379 distinct token pairs after the change, 378 after the repair. Score after repair: 0.77. The original passage was restored exactly.
Changed the ‘a’ in ‘quality’ to a Cyrillic letter that looks the same.
Sign the sheet, state the ambient conditions, and send the signed result to the quаlity file.
Can the original text be restored?
Yes, for this passage. Making the same replacement in genuine Cyrillic text would be unsafe.
Technical detail
Replaced Latin Small Letter A, U+0061 with Cyrillic Small Letter A, U+0430 in the word “quality” at position 1437.
The normaliser reversed the allow-listed character change using the submitted text alone.
Scroll horizontally to see all columns
| Position | Submitted value | Normalised value | Reason |
|---|---|---|---|
| 1437 | Cyrillic Small Letter A, U+0430 | Latin Small Letter A, U+0061 | Allowed mapping for this passage |
379 distinct token pairs after the change, 378 after the repair. Score after repair: 0.77. The original passage was restored exactly.
Deleted the ‘c’ in ‘clock’.
Write down the meter value, the ambient temperature, and the lock time.
Can the original text be restored?
No. The altered text gives us no reliable way to recover the missing letter.
Technical detail
Deleted Latin Small Letter C, U+0063 from the word “clock” at position 231. One visible character is now absent.
The missing character is not present in the transformed text. The normaliser does not guess missing text.
Scroll horizontally to see all columns
| Position | Submitted value | Normalised value | Reason |
|---|---|---|---|
| 231 | Absent character | No change | No repair rule for deleted text |
378 distinct token pairs after the change, 378 after the repair. Score after repair: 0.65. The original passage was not restored.
Every tenth word was targeted. Of those 42 words, 39 contained a letter that could be swapped for a lookalike. The editing rate was 0.09, close to the rate used in published attack research.
Set the reference flow to 40 litres per minute before yоu record the first value on the raw sheet. Hold thе line at that flow for two minutes so the рressure can settle.
Can the original text be restored?
Yes, for this passage. Every edit uses a listed lookalike, so all 39 changes are reversed exactly.
Technical detail
Replaced one Latin letter with its Cyrillic lookalike in 39 of the 427 scored words, an editing rate of 0.09. Every tenth word was targeted, 42 in all, and 3 of those held no letter on the swap list, so they were left alone. The passage still reads normally.
The normaliser reversed the allow-listed character change using the submitted text alone.
Scroll horizontally to see all columns
| Position | Submitted value | Normalised value | Reason |
|---|---|---|---|
| many | 39 Cyrillic lookalikes | All mapped back to Latin | Allowed mappings for this passage |
408 distinct token pairs after the change, 378 after the repair. Score after repair: 0.77. The original passage was restored exactly.
These repairs cover only the prepared changes shown here. Applying them to arbitrary text could damage another language.
A rewrite of the same passage
The rewritten passage is shorter, and it still walks through the same flow-meter check. Its different phrasing moves the score from 0.77 to -1.32, giving a result of below threshold.
The scorer sees the changed words regardless of who rewrote them. This result describes one registered rewrite, and other revisions can fall on either side of the cut-off. There is nothing for character repair to reverse here. The score says nothing about who rewrote the text, or about whether a model had a hand in it earlier.
Each score uses the full prepared passage. The cards show their opening lines.
- Full-passage result
- Below threshold
- Full-passage score
- 0.77
Set the reference flow to 40 litres per minute before you record the first value on the raw sheet. Hold the line at that flow for two minutes so the pressure can settle.
- Full-passage result
- Below threshold
- Full-passage score
- -1.32
Begin at a reference flow of 40 litres per minute and write the first reading only after the line has held that rate for two minutes, long enough for the pressure to come to rest.
Testing older writing
6 of 96
6 passages crossed the cut-off by chance. All 96 passages were written long before modern language models existed. These are the full-passage results. The method also reports the registered 199, 200 and 201-pair prefixes.
All observers testify to the prodigious volume of voice possessed by these animals. According to the writer whom I have just cited, in one of them, the Siamang, "the voice is...
Thomas Henry Huxley, Evidence as to Man's Place in Nature
Score 2.91
The evil is perhaps gone too far to be remedied, but I feel little doubt in my own mind that if the poor laws had never existed, though there might have been a few more instances...
T. R. Malthus, An Essay on the Principle of Population
Score 2.43
The Dutch and French colonies, though under the government of exclusive companies of merchants, which, as Dr Adam Smith says very justly, is the worst of all possible governments,...
T. R. Malthus, An Essay on the Principle of Population
Score 2.42
Attention and repetition help much to the fixing any ideas in the memory. But those which naturally at first make the deepest and most lasting impressions, are those which are...
John Locke, An Essay Concerning Human Understanding
Score 2.60
Thus many of those ideas which were produced in the minds of children, in the beginning of their sensation, (some of which perhaps, as of some pleasures and pains, were before...
John Locke, An Essay Concerning Human Understanding
Score 2.30
Beside the improvements in arts and machinery, there are various other causes which are constantly operating on the natural course of trade, and which interfere with the...
David Ricardo, Principles of Political Economy
Score 2.49
How was this measured?
| Evaluation set | 96 passages by 8 authors |
|---|---|
| Threshold | 2.30 |
| Measured rate | 6.25% (6 of 96) |
| Wilson 95% interval | 2.9% to 12.97% |
| Calibration target | 2% |
With only 96 passages, 6 crossings give a wide interval. Several passages came from each author, so do not read it as a population estimate for human writing.
What to write down while working
Six months on, someone asks how a piece of work was made. The finished file cannot answer them; most of what happened never reached it. So write things down while they happen. Where the material came from. What the model produced. What a person corrected or checked, who signed it off, where it went afterwards. Dull notes, mostly, and also the only place an explanation can come from once the result is questioned.
Forcing the final text into an “AI” or “human” box throws all of that away.
Source of the material
Note the source, version, date and licence. Link to it when you can; if it is sensitive, point to a secure record.
The model's part, if any
Note where it drafted, summarised, translated, calculated or changed something. Keep only the prompts needed to explain those changes, and leave private material out.
Checks and changes made by a person
Note what the person checked, changed, rejected or tested, including sources they checked without the model. They should be able to explain the finished work.
Review, approval and later use
Note who reviewed or approved it. Keep the versions that mattered, and any later corrections or reuse.
You need enough of a record to explain the work, not a log of every keystroke. Sensitive material should have an owner, access rules and a date it gets deleted. A missing record is a gap to investigate; it is not proof of misconduct.
Part of that history can travel in a signed record. Only part. The signature does no judging either; a person still has to read what was asserted and decide how far to trust it. Read about signed records.
Ask for the checks, not a score
When you need to know how a piece of work was made, ask the person who made it to show the checks they ran, the failures, and what they threw out, then to explain why the conclusion follows from what is left. Someone who can reproduce the work and talk through those choices tells you more than any detector score.
Related training
The first engineering module is called Verification Before Trust. This page applies the same rule to text watermarks, and the wider course is about using AI without lowering the standard of the work.
Both courses run as four live sessions of forty minutes, weekly, online at United Kingdom times, taught by the person who wrote them. You bring whichever assistant you already use: Copilot, ChatGPT, Gemini or Claude. Both cost £495 including VAT per place.
The engineering course
For people who write and review code, from two months to two decades in.
Prepared repositories with real ambiguity and real risk in them, reviewed by the tutor. You keep the change you would merge.
The business course
For analysts, marketing, operations, HR, finance and product. No coding needed.
Document packs with problems planted in them, reviewed by the tutor. You keep the corrected briefing, the written definition and the customer pack you produce.
Teams can run either course privately, usually six to fifteen people, at the same price per place. Talk to us about a private cohort, or see everything on the training site.