A byte-level scan of collected Claude output
Published ; updated
The Declawd measurements in this article use archived v1. Their original figures and evidence are preserved. The current method and interactive pages use v2, so links from this article to them show v2 figures.
The hidden-character theory can be tested, so we tested it. We collected 120 outputs assigned to four Claude models, 457,045 characters in all, and inspected every non-ASCII character in them. None belonged to a hidden class that this collection route could preserve and the scanner could reliably observe. Anthropic later said its text watermark adds no hidden characters. The scan tested that discarded carrier theory. It did not test SynthID-Text.
Corpus
We used 30 fixed prompts for each of four model assignments, producing 120 files. The prompts covered factual prose, informal posts, technical writing, lists, Python, JSON, dialogue, accented text and instructions to use ASCII. Every model received the same prompt bytes.
A separate worker wrote the files through its file-writing tool. They did not pass through a chat renderer, terminal or clipboard. The files give us a close view of one collection path, but not the raw Anthropic API response.
Two limits turned up. The positive control asked for six non-ASCII space variants and all six reached disk as ordinary spaces, so the route could not validate that class at all. One Fable file also carries a disclosed phrase change made after generation. And the model names? They come from the collection setup and the worker reports. There was no API response header to verify them against.
Scanner coverage
The scanner listed every character above ASCII. Zero-width spaces and joiners were checked, along with the word joiner, byte-order marks and soft hyphens. So were the invisible mathematical operators, the line and paragraph separators and combining marks. Bidirectional controls and isolates got attention too, as did variation selectors and Unicode tag characters. Then the lookalikes. It hunted for Cyrillic and Greek letters sitting inside Latin words, and for three apostrophe lookalikes that a separate Claude Code investigation had turned up.
Disk
Scroll horizontally to see all columns
| Model assignment | Characters | Detectable hidden class | Other non-ASCII |
|---|---|---|---|
| Opus 5 | 125,943 | 0 | 15 em dashes, accented letters |
| Sonnet 5 | 107,135 | 0 | 3 em dashes, accented letters |
| Haiku 4.5 | 113,705 | 0 | 7 em dashes, 6 check marks, accented letters |
| Fable 5 | 110,262 | 0 | 12 em dashes, accented letters |
No zero-width character, bidirectional control, tag character or variation selector reached disk in any of the 120 files. All three apostrophe lookalikes were absent too.
The remaining characters have ordinary uses. Thirty-six of the 37 em dashes occur in dialogue prompts, where they mark interrupted speech. The remaining one is in an informal post. Six check marks appear inside printed test output in generated Python. Accented letters appear mainly in the 12 prompts that explicitly asked for correct diacritics, with a few elsewhere. All 12 ASCII tasks contained only printable ASCII and line breaks.
These counts apply only to this corpus and route. They are not a general estimate of how often a hidden mark appears in Claude output.
Collection
A clean scan means little if the collection route stripped the characters first. We asked a model to emit known hidden characters through the same route.
Zero-width characters, bidirectional controls, variation selectors, combining marks and Cyrillic lookalikes reached disk intact. A Unicode tag payload also survived and still decoded to the text placed inside it. The scanner detected those tested classes when they were present.
The space test failed. All six requested non-ASCII spaces reached disk as ordinary spaces, so we claim nothing about a watermark carried only by space variants. Was the filesystem stripping them? No. A non-breaking space written by Python came through the filesystem, the clipboard and JSON untouched.
Humans
We also scanned 3,059,329 characters from five Project Gutenberg novels, downloaded as raw files.
Scroll horizontally to see all columns
| Corpus | Characters | Curly quote and apostrophe characters | Bidirectional controls |
|---|---|---|---|
| Collected Claude output | 457,045 | 0 | 0 |
| Five nineteenth-century novels | 3,059,329 | 20,224 | 2 |
The two bidirectional controls in the human corpus sit around Hebrew text in the Project Gutenberg file for Moby-Dick. They are modern digital direction markers in a transcription of an 1851 book. Any rule treating these characters as proof of machine writing would flag the human file and clear all 120 collected outputs.
Recheck
Reported result, checked 24 September 2026. A separate scan published on 12 August covers 7.2 million characters and reports no anomalous hidden Unicode or whitespace encoding. Worth reading with two caveats in hand. The post does not document a class-for-class comparison with this site's registry, and the response corpus behind the numbers is not published with it.
Reported result, checked 24 September 2026. The same author's green/red-list constrained-choice permutation tests were negative (one permutation p-value, on Fable 5, came out at 0.886). Separate SynthID-style probes returned strong Z(H) prompt effects, but no clean context-window boundary identifying a watermark. The author interprets those effects as inconclusive for SynthID. An addendum dated 20 August reports serving behaviour changing without a newly identifiable watermark signature. That is not a finding that the tested models are unmarked.
Reported result, checked 24 September 2026. A later black-box study at revision 8838e0d builds on the same Red-Green test and changes how the test number is shown. With the number given as spaced digits, so that the phrase the watermark reads first appears in the response, the study reports a keyed pattern in Fable 5.1 output tested on 2 September: a corrected p-value of 0.006 with a 15-digit number, replicated, and close to zero with a 40-digit one. It reports Sonnet 5 as null on the same prompts, which fits Anthropic's table, and Opus 4.8 with a small trace whose per-digit pattern does not match Fable's. This asks a different question from the scan above. It tests whether a model's sampling carries a keyed bias across many generated answers, so it cannot say whether any one passage is marked. The p-values are the author's, and Declawd has not rerun the study.
Limits
120 files, 457,045 characters, four model assignments, one collection path, one day. Within that corpus, not a single character from the detectable hidden classes reached disk. A cleaner pointed at those characters would have had nothing to remove.
Every assigned model predates Anthropic's 2 August launch rule. The scan did not reach the raw API body, claude.ai, the desktop app or a Copy button. The route could not preserve the six space variants, and we could not verify model identity independently of the collection setup.
A byte scan cannot see a watermark carried by token selection. The marked passage on this site scores 2.99 against a cut-off of 1.80 and still contains ordinary ASCII. Its pattern appears only when token pairs are scored with the published seed. The scan tells us which characters are present. It cannot tell us how the words were selected.
Official statement, published 14 August and checked 24 September 2026. Anthropic identifies Claude's text mark as a version of SynthID-Text and says nothing is added to the text. The original null result is compatible with that description. It does not validate the production system, and every model assignment in this corpus pre-dates the 2 August launch rule. By 24 September Anthropic's Help Centre listed one of them, Opus 5, with text watermarks, and Sonnet 5, Haiku 4.5 and Fable 5 with C2PA credentials only (their text is due by 2 December). The corpus was generated on 12 August, before the announcement, so the null result describes those files and says nothing about how the models are served now.
Looking
Reported result, checked 24 September 2026. One class on that scanner list deserves a second look. aloshdenny/claude-awm at revision 90f3975 reports that one family of edits, inserting variation selectors, pushes its open SynthID-Text detector below threshold while ordinary normalisation leaves the inserted selectors in place. That revision reports seven open models, 440 main study cells and a separate scaling sweep to one million tokens. Its required insertion rate is not fixed. On gpt-oss-20b it reports the rate rising from 9.3 per cent at 1,024 tokens to 25 per cent at 131,072 tokens. Across the longer sweep, all 45 measured cells at 35 per cent insertion or above fell below its threshold. The repository also marks a flat Qwen long-context baseline as unreliable in absolute terms. This is a length-dependent result for one edit family, not a claim that one inserted character defeats a detector.
This is a different question from the one our scan answered. We looked for these characters sitting in unaltered output and found none. The study inserts them on purpose as a carrier against a statistical detector. A scan for incidental presence and an adversarial edit are not the same test.
The mechanism is worth stating exactly, because it is often stated wrongly. Normalisation does not fold away a zero-width space, and it does not fold away a variation selector either. What separates them is what a cleaner chooses to remove: it strips zero-width and direction characters as a matter of policy, and it leaves variation selectors alone because they carry real meaning in emoji and in Chinese, Japanese and Korean text. A character a cleaner is right to preserve is, for that same reason, one a character cleaner cannot be relied on to remove.
None of this is a result about Claude. The study runs an untrained detector on open models, not Anthropic's production system, and says so. We have not reproduced it.
The low-entropy point has separate support. A paper accepted at COLM 2026 reports that existing token-level watermarking is inadequate on constrained output such as translation, summarisation and code, where there are few word choices to carry a mark, and proposes a sequence-level method it says recovers up to 28 per cent in detection F1. That is a limit of today's token-level schemes, not of watermarking as such, and it too is measured on open schemes rather than Claude. It lines up with Anthropic's own note that exact code offers few eligible choices.
Sampling
How its own sampling worked, or what happened to the text after generation, is not something a language model can reliably see. So what do you get if you ask one whether its answer is watermarked? A plausible account. Not a measurement.
Inspect the output instead. A proper Unicode check lists every code point above U+007F with its name and its position. One more step before you trust a clean result: push known hidden characters through the same route and confirm they both survive and get detected.
The corpus, controls, scanner, tests and SHA-256 manifest are public in san-digital/claude-watermark-audit. The repository records 177 artefacts and the exact bytes used for every count.
Read the punctuation count, or inspect the watermark method used on this site.