Table 1: Synoptic table of the families of text watermarks. The last column is the one that matters for this article: in every family except the Unigram scheme, an edit damages not only the position it touches but every position that used it as seeding context. It is this asymmetry which the geometry of Section 4 and the identity of Section 5 make precise.
Family
Mechanism
Context h
Output law
What an edit destroys
Green list [7]
logit bias
≥1
perturbed
token and its h predecessors
Unigram [15]
logit bias
0
perturbed
the token alone
Exponential [10, 2]
sampling rule
≥0
exact
token and seeding context
Tournament [5]
sampling rule
≥1
exact
token and seeding context
Semantic [6]
sentence partition
—
perturbed
the sentence embedding
Table 2: This table summarizes the independence of the endpoint and path data, in dimension n=64. All seven chains share the same first and last state, so the semantic deficit is constant to the last recorded digit; the holonomy energy is not. Note also that the signature, in the last column, moves in steps and is constant on pairs of rows, exactly as Proposition 3.1 predicts it must.
detours
δ𝒰
η(𝒰)
sign(In−R𝒰∗)
0
0.12241744
0.0000
(2,0,62)
1
0.12241744
0.2454
(2,0,62)
2
0.12241744
0.5190
(4,0,60)
3
0.12241744
0.6898
(4,0,60)
4
0.12241744
0.8260
(6,0,58)
5
0.12241744
0.9427
(6,0,58)
6
0.12241744
1.0465
(8,0,56)
Table 3: Residual detector statistic, as a fraction of the original, for the green-list scheme with h=1 under three edit patterns of identical retention rate. Counting the intact-window set directly from each edit pattern, Theorem 5.1 predicts 0.947,0.897,0.797,0.697,0.597,0.496 for the middle column and 0.902,0.802,0.602,0.401,0.201,0.000 for the right-hand one. The largest discrepancy is 0.008 and the typical one 0.003, over 120 sequences per cell; and the periodic pattern at ρ=0.5 is predicted to give exactly zero, and does.
ρ
independent
contiguous block
periodic
0.95
0.904
0.947
0.903
0.90
0.806
0.897
0.802
0.80
0.637
0.795
0.595
0.70
0.492
0.697
0.399
0.60
0.362
0.594
0.196
0.50
0.247
0.488
−0.005
Table 4: The six chains; medians over the ninety passages of the three schemes, except the last column, which counts the chains still detected above z=4. Ordered by semantic deficit, every column moves monotonically: the meaning drifts, the path lengthens, the surface is retained less and the mark fades, all together. That is precisely why the endpoint alone cannot be read as a measure of attack strength, and why the test below holds it fixed.
chain
L
δ𝒰
η(𝒰)
ρ
|I|/T′
residual
detected
Spanish
2
0.064
0.083
0.767
0.679
0.680
84/90
German
2
0.071
0.092
0.733
0.651
0.633
83/90
French
2
0.073
0.094
0.722
0.616
0.625
82/90
German twice
4
0.082
0.101
0.692
0.598
0.591
82/90
German, French
4
0.109
0.118
0.644
0.531
0.554
67/90
French, German, Spanish
6
0.116
0.129
0.594
0.501
0.496
66/90
Table 5: Theorem 5.1 against the 538 chains whose statistic is finite; medians. The measured intact-window fraction predicts the median residual to within one hundredth for the context-free scheme, three for the green-list scheme and seven for the exponential one. The independent-edit corollary ρh+1, which needs no measurement of the attacked text at all, happens to fall closer for the green-list scheme and much further for the exponential one, where it is off by twelve hundredths; and chain by chain it is the measured fraction that follows the residual, correlating +0.67 with it against +0.54 for the retention rate. The median absolute error per chain is 0.083, 0.054 and 0.088, which the length correction of (6.1) moves to 0.086, 0.067 and 0.081: at this depth of translation the length is preserved in the median, and the correction has little to do.
scheme
h
ρ
ρh+1
|I|/T′
Tatt′/T0′
(6.1)
observed
green-list
1
0.700
0.490
0.542
0.972
0.553
0.509
unigram
0
0.694
0.694
0.694
0.972
0.708
0.685
exponential
1
0.692
0.478
0.531
0.978
0.535
0.601
Table 6: The preregistered regression. The predictors are standardised, so that the coefficients may be compared; the corrected threshold is 3.3×10−3. The last column is a robustness check which the clause did not ask for: the same coefficient with a standard error clustered on the base passages, of which there are only thirty, so that the check is a severe one. Under it the unigram scheme still clears the threshold and the green-list scheme no longer does.
scheme
n
R2
βρ
βδ
βη
pη
clustered
green-list
180
0.41
+0.051
−0.023
−0.070
1.7×10−3
3.5×10−2
unigram
180
0.49
+0.064
+0.278
−0.145
4.0×10−9
8.2×10−4
exponential
178
0.34
+0.116
−0.060
−0.014
0.74
0.77
Table 7: Every numerical claim of Section 6 and the run that produced it. Directory names are the script name prefixed by run_ and suffixed by the stamp of the third column, under 7. Results/Article_LLW/. The corpus was built in three successive invocations, each carrying forward the chains of the one before, so that an interruption on a machine of this size would cost at most one of them; the stamp given is that of the last, whose manifest records the provenance of the other two. The final row is the targeted re-audit of the records which had failed to regenerate, discussed below.
Statistical watermarks for language models live in the freedom of the signifier: they choose among tokens that are nearly equivalent in meaning, and they are therefore eroded by exactly those transformations which move the form of a text while leaving its content in place. The literature measures such transformations by their endpoint, through the semantic similarity between the original and the rewritten text. We show that the endpoint is the wrong statistic. Adapting the formalism of linguistic loops, we prove that the invariant of a chain of meaning-preserving transformations factorises canonically into an endpoint part and a holonomy in the stabiliser of the initial state, the second of which the semantic deficit cannot see; the loop rotation is parallel transport on the unit sphere of the embedding space, so that the analogy with the Wilson loop becomes a theorem rather than a figure of speech. On the side of the detector we prove an exact identity: the residual statistic is proportional to the number of positions whose seeding window survived intact, from which the decay law $\rho^{h+1}$ follows as the independent-edit corollary. The identity has a disconcerting consequence, which we confirm to three decimal places: at one and the same retention rate the surviving signal may be one half of the original, one quarter of it, or exactly nothing, according only to where the edits fall.