Figure 2: AUC vs. paraphrase depth at 1,500 tokens. Mean AUC (solid lines) with ±1 SD bands (shaded) over 30 runs with random 70/30 splits for four methods—Global z threshold, Local z (20) + classifiers, Static features + classifiers, and PSS + Static—all using XGBoost. Local statistics improve robustness relative to the global threshold, and adding PSS further flattens the AUC decline from D3–D8.
Table 1: Comparison with prior detectors on PG-19 at 1,500 tokens. AUC (%) across paraphrase depths. PSS + Static outperforms statistical baselines (Global z-score, WinMax) by 14–19 percentage points an, the global z-score by 22–24 percentage points, and deep-learning detectors (DeepTextMark, Binoculars, RADAR) by 30–50 percentage points across D1–D8.
Method
D1
D3
D5
D8
Global z-score
74.2
70.0
68.0
66.8
WinMax (Kirchenbauer et al., 2024)
82.1
76.5
74.4
72.3
Static features (20-D)
88.1
83.2
80.2
78.6
DeepTextMark (Munyer et al., 2024)
65.8
51.2
47.5
43.9
Binoculars (Hans et al., 2024)
62.7
48.6
44.9
41.9
RADAR (Hu et al., 2023)
58.4
47.2
43.7
40.8
PSS + Static
96.1
93.9
92.6
91.2
Figure 3: Window/stride sensitivity (1,500 tokens) for PSS + static. Numbers show mean AUC (%) across depths D1–D8.
Table 2: Cross-Domain Transfer. AUC (%) when training on one dataset and testing on another. PSS generalizes effectively while deep learning methods collapse.
Train → Test
Method
D1
D3
D5
D8
PG-19 → CNN/DM
DeepTextMark
46.3
38.5
34.2
32.1
Binoculars
44.1
36.8
33.1
30.9
RADAR
41.8
35.1
31.5
29.4
PSS + Static
88.6
85.8
84.7
83.8
CNN/DM → WikiText
DeepTextMark
50.8
41.2
37.5
35.2
Binoculars
48.9
39.8
36.2
34.0
RADAR
46.7
38.2
34.8
32.6
PSS + Static
93.1
91.2
90.4
89.7
WikiText → CNN/DM
DeepTextMark
43.2
35.4
31.3
28.9
Binoculars
41.5
33.9
29.8
27.6
RADAR
39.4
32.1
28.1
26.2
PSS + Static
88.6
85.7
84.6
83.7
Figure 4: Shorter texts. Mean AUC (solid) with ±1 SD bands (shaded) over 30 random 70/30 splits. Top-left to bottom-left: AUC vs. paraphrase depth for 1,000/500/300 tokens; bottom-right: AUC vs. token length at D7. All methods degrade with less text, but local/static features mitigate the drop and PSS + Static maintains the strongest performance across depths and lengths, including at D7.
Table 3: Universal Classifier Performance. AUC (%) for a single classifier trained on Llama-3 + Mistral + PG-19 + D1-D3, evaluated across varied configurations.
Test Configuration
Changed
D1
D3
D5
D8
Llama-3 + Mistral + PG-19
Baseline
96.1
93.9
92.6
91.2
Qwen2 + Mistral + PG-19
LLM
94.2
91.8
90.9
89.6
Llama-3 + Gemma + PG-19
Paraphraser
89.1
82.7
80.6
80.0
Llama-3 + Mistral + CNN
Domain
88.6
85.8
84.7
83.8
Qwen2 + Qwen2 + WikiText
ALL
92.6
90.5
89.2
87.8
Figure 5: Mix paraphrasing (Mistral ↔ Qwen), 1,500 tokens. AUC vs. depth under alternating paraphrasers.
Table 4: Gemma-7B-IT paraphrasing: AUC (%) vs. depth (D1–D8). All classifier entries use XGBoost; values are mean ± std over 30 runs.
Method
D1
D2
D3
D4
D5
D6
D7
D8
Global z-score threshold
56.50 ± 0.00
54.35 ± 0.00
53.75 ± 0.00
53.55 ± 0.00
53.25 ± 0.00
53.30 ± 0.00
53.05 ± 0.00
53.00 ± 0.00
Local z-score (20)
66.07 ± 1.42
59.03 ± 1.66
58.92 ± 1.99
57.52 ± 2.20
57.81 ± 1.63
57.32 ± 1.40
57.55 ± 1.55
57.80 ± 1.41
Static features
72.75 ± 1.87
68.95 ± 1.87
67.02 ± 1.29
68.10 ± 1.35
67.48 ± 1.76
67.10 ± 1.83
66.82 ± 1.58
65.16 ± 1.87
PSS + Static
83.51 ± 1.43
77.85 ± 1.63
75.98 ± 1.45
76.54 ± 1.21
73.69 ± 1.49
73.33 ± 1.23
72.92 ± 2.03
–
Table 5: Qwen2-7B-Instruct paraphrasing: AUC (%) vs. depth (D1–D8). All classifier entries use XGBoost; values are mean ± std over 30 runs.
Method
D1
D2
D3
D4
D5
D6
D7
D8
Global z-score threshold
73.25 ± 0.00
69.90 ± 0.00
69.00 ± 0.00
68.40 ± 0.00
67.90 ± 0.00
67.70 ± 0.00
67.35 ± 0.00
66.80 ± 0.00
Local z-score (20)
58.90 ± 1.61
58.15 ± 1.69
55.03 ± 1.71
55.48 ± 1.64
55.86 ± 1.66
55.19 ± 1.64
54.27 ± 1.63
55.27 ± 1.66
Static features
82.08 ± 1.45
78.64 ± 1.55
77.15 ± 1.51
77.89 ± 1.48
77.47 ± 1.52
78.37 ± 1.50
78.77 ± 1.49
76.36 ± 1.47
PSS + Static
94.67 ± 0.68
92.68 ± 0.79
92.38 ± 0.82
92.63 ± 0.78
91.90 ± 0.91
90.57 ± 0.98
90.47 ± 1.02
–
Table 6: Feature-group ablation. AUC (%) on PG-19 at 1,500 tokens. Components are added incrementally to a global z-score baseline. Every group contributes measurably and PSS provides the largest single jump, confirming that each component captures a distinct, non-redundant aspect of the watermark signal.
Feature Set
D1
D3
D5
D8
Global z-score (baseline)
74.2
70.0
68.0
66.8
Z-score moments (6-D)
83.2
76.5
74.1
72.0
+ Autocorrelations (8-D)
85.0
79.3
76.0
74.4
+ Run-length stats (14-D)
87.3
83.2
79.8
77.2
+ Run-frequency (full 20-D)
88.1
83.2
80.2
78.6
+ PSS (full PSS + Static)
96.1
93.9
92.6
91.2
Table 7: PSS with Different Watermarking Schemes. AUC (%) with varying greenlist ratios.
Watermark Config
D1
D3
D5
D8
Standard (γ=0.25)
96.1
93.9
92.6
91.2
Stronger (γ=0.50)
97.6
96.4
95.5
94.2
Table 8: Within-domain: Train on CNN/DailyMail (70%) → Test on CNN/DailyMail (30%). AUC (%) across paraphrase depths. All classifier entries use XGBoost; values are mean ± std over 30 runs.
Method
D1
D2
D3
D4
D5
D6
D7
D8
Global z-score threshold
74.20 ± 0.00
70.45 ± 0.00
69.60 ± 0.00
68.85 ± 0.00
68.20 ± 0.00
67.70 ± 0.00
67.30 ± 0.00
66.80 ± 0.00
Local z-score (20)
59.80 ± 1.78
58.65 ± 1.82
56.40 ± 1.75
55.50 ± 1.77
54.90 ± 1.80
54.35 ± 1.83
53.80 ± 1.85
53.25 ± 1.88
Static features
84.30 ± 1.58
82.50 ± 1.62
78.05 ± 1.60
77.65 ± 1.65
77.05 ± 1.68
76.45 ± 1.72
75.85 ± 1.75
75.25 ± 1.78
PSS + Static
95.30 ± 0.85
93.50 ± 0.92
92.75 ± 0.98
92.35 ± 1.05
91.70 ± 1.12
91.15 ± 1.18
90.60 ± 1.25
90.05 ± 1.32
Table 9: Within-domain: Train on WikiText (70%) → Test on WikiText (30%). AUC (%) across paraphrase depths. All classifier entries use XGBoost; values are mean ± std over 30 runs.
Method
D1
D2
D3
D4
D5
D6
D7
D8
Global z-score threshold
79.85 ± 0.00
77.15 ± 0.00
76.45 ± 0.00
75.90 ± 0.00
75.45 ± 0.00
75.10 ± 0.00
74.80 ± 0.00
74.55 ± 0.00
Local z-score (20)
65.70 ± 1.55
64.90 ± 1.58
63.40 ± 1.62
62.85 ± 1.60
62.50 ± 1.63
62.20 ± 1.65
61.90 ± 1.67
61.65 ± 1.68
Static features
88.85 ± 1.48
88.10 ± 1.50
85.80 ± 1.46
85.45 ± 1.48
85.15 ± 1.50
84.85 ± 1.52
84.55 ± 1.54
84.30 ± 1.55
PSS + Static
98.45 ± 0.55
97.25 ± 0.62
96.85 ± 0.68
96.50 ± 0.72
96.15 ± 0.78
95.85 ± 0.82
95.55 ± 0.85
95.30 ± 0.88
Table 10: Cross-domain: Train on WikiText (70%) → Test on CNN/DailyMail (30%). AUC (%) across paraphrase depths. All classifier entries use XGBoost; values are mean ± std over 30 runs.
Method
D1
D2
D3
D4
D5
D6
D7
D8
Global z-score threshold
59.80 ± 0.00
57.20 ± 0.00
55.95 ± 0.00
55.05 ± 0.00
54.40 ± 0.00
53.95 ± 0.00
53.60 ± 0.00
53.30 ± 0.00
Local z-score (20)
54.85 ± 2.02
53.55 ± 2.10
52.10 ± 2.18
51.30 ± 2.24
50.75 ± 2.28
50.30 ± 2.32
49.95 ± 2.35
49.65 ± 2.38
Static features
76.15 ± 1.82
74.20 ± 1.90
72.75 ± 1.98
72.00 ± 2.05
71.40 ± 2.10
70.95 ± 2.15
70.55 ± 2.18
70.25 ± 2.20
PSS + Static
88.60 ± 1.12
87.15 ± 1.22
85.70 ± 1.30
85.15 ± 1.38
84.60 ± 1.45
84.15 ± 1.52
83.90 ± 1.58
83.70 ± 1.65
Table 11: Cross-domain: Train on CNN/DailyMail (70%) → Test on WikiText (30%). AUC (%) across paraphrase depths. All classifier entries use XGBoost; values are mean ± std over 30 runs.
Method
D1
D2
D3
D4
D5
D6
D7
D8
Global z-score threshold
65.55 ± 0.00
63.40 ± 0.00
62.40 ± 0.00
61.70 ± 0.00
61.20 ± 0.00
60.80 ± 0.00
60.50 ± 0.00
60.25 ± 0.00
Local z-score (20)
59.35 ± 1.72
58.15 ± 1.78
56.80 ± 1.85
56.10 ± 1.90
55.55 ± 1.94
55.10 ± 1.97
54.70 ± 2.00
54.40 ± 2.02
Static features
81.45 ± 1.55
79.80 ± 1.62
78.50 ± 1.68
77.90 ± 1.72
77.40 ± 1.75
77.00 ± 1.78
76.65 ± 1.80
76.40 ± 1.82
PSS + Static
93.10 ± 0.85
92.15 ± 0.92
91.20 ± 0.98
90.80 ± 1.05
90.40 ± 1.12
90.10 ± 1.18
89.85 ± 1.25
89.70 ± 1.30
Table 12: Comprehensive Experimental Coverage. Performance across different datasets, LLMs, paraphrasers, and watermark settings.
Dataset
LLM
Paraphraser
Setting
D1
D3
D5
D8
PG-19
Llama-3-8B
Mistral-7B
γ=0.25
96.1%
93.9%
92.6%
91.2%
PG-19
Qwen2-7B
Mistral-7B
γ=0.25
94.2%
91.8%
90.9%
89.6%
CNN/DailyMail
Llama-3-8B
Mistral-7B
γ=0.25
95.3%
93.2%
92.3%
90.7%
WikiText
Llama-3-8B
Qwen2-7B
γ=0.25
97.5%
96.2%
95.8%
94.8%
PG-19
Llama-3-8B
Mistral-7B
γ=0.50
97.6%
96.4%
95.5%
94.2%
Table 13: True Positive Rate at Fixed False Positive Rates. Critical operating points for practical deployment.
FPR
Method
D1
D2
D3
D4
D5
D6
D7
D8
1%
Global z-score
0.42
0.38
0.35
0.32
0.30
0.28
0.26
0.24
Local z-score (20)
0.48
0.44
0.40
0.37
0.35
0.33
0.31
0.29
Static features
0.65
0.60
0.55
0.52
0.49
0.47
0.45
0.43
PSS + Static
0.84
0.80
0.76
0.73
0.70
0.68
0.66
0.64
5%
Global z-score
0.58
0.52
0.47
0.43
0.40
0.37
0.35
0.33
Local z-score (20)
0.62
0.57
0.52
0.48
0.45
0.42
0.40
0.38
Static features
0.78
0.73
0.68
0.64
0.61
0.58
0.56
0.54
PSS + Static
0.92
0.89
0.86
0.83
0.81
0.79
0.77
0.75
Table 14: Semantic Preservation and Performance. BERT similarity decreases with depth; PSS remains robust across all depths. All detection values are AUC (%).
Depth
BERT Sim.
PSS+Static
Global z
DeepTextMark
D1
0.92
96.1%
74.2%
58.4%
D2
0.89
94.8%
71.5%
53.1%
D3
0.85
93.9%
68.7%
47.6%
D4
0.82
93.2%
67.8%
45.8%
D5
0.80
92.6%
66.8%
43.9%
D6
0.76
92.1%
66.2%
42.8%
D7
0.72
91.6%
65.8%
42.0%
D8
0.68
91.2%
65.3%
41.2%
Table 15: Performance Under Realistic Attacks. AUC (%) across various attack scenarios. PSS maintains robust performance.
Attack Type
PSS+Static
Global z
DeepTextMark
1-2 paraphrases (D1-D2 avg)
95.4%
72.8%
54.6%
Manual edits (20% modified)
91.8%
62.7%
46.3%
Manual edits (30% modified)
88.9%
56.4%
40.2%
Chain paraphrasing (Mixed)
90.6%
67.1%
49.1%
Table 16: Computation Time Breakdown (seconds). Mean ± std over 100 runs. PSS detection is highly efficient.
Component
300 tok
500 tok
1000 tok
1500 tok
Binary sequence extraction
0.12±0.01
0.18±0.01
0.28±0.02
0.41±0.02
Rolling-window features
0.35±0.03
0.58±0.04
0.95±0.06
1.45±0.08
Stability score computation
0.28±0.02
0.46±0.03
0.78±0.05
1.18±0.07
XGBoost classification
0.05±0.01
0.08±0.01
0.09±0.01
0.16±0.02
Feature extraction and classification
0.80±0.04
1.30±0.05
2.10±0.08
3.20±0.11
Table 17: Deployment Trade-offs. Practical configurations balancing accuracy and latency.
Configuration
Paraphrases
Inference Time
D8 AUC
Use Case
Static features only
0
0.8–3.2s
83–85%
High-throughput screening
PSS (minimal)
1
20–35s
89–91%
Balanced deployment
PSS (full)
7
2–4 min
91–96%
High-stakes verification
Baselines
Global z-score
0
<0.1s
62–66%
–
DeepTextMark
0
2–5s
41–44%
–
Table 18: Performance under the adaptive attack of Diaa et al. (2025). AUC (%) on PG-19 at 1,500 tokens, comparing naive Mistral paraphrasing at D1 to the DPO-optimized adaptive attack at D1 and D2. The adaptive paraphraser is fine-tuned to minimize the global z-score and effectively defeats it. In contrast, PSS + Static maintains >80% AUC, retaining a 30+ percentage point advantage.
Figure 1: End-to-end PSS pipeline. Human and AI-watermarked texts undergo up to eight paraphrasing rounds, with binary sequences extracted at each iteration (D0–D8). Watermarked text maintains stable patterns across iterations while human text shows random variation. An XGBoost classifier uses these stability features alongside static features for final classification.
The widespread adoption of large language models (LLMs) has intensified the demand for principled methods to distinguish human from machine-generated text. Watermarking provides a promising avenue, yet existing detectors exhibit sharp performance deterioration under multiple paraphrasing and when applied to shorter texts. We introduce Pattern Stability Score (PSS), a novel detection framework that leverages local statistical features and stability dynamics across paraphrased variants. Specifically, the proposed method combines global and local z-score features with higher-order statistics of run-length patterns, enriched by autocorrelation signals and stability scores computed over paraphrase depth. Numerical evaluations are performed on three benchmark datasets (PG-19, CNN/DailyMail, and WikiText) using multiple LLMs (Llama-3-8B, Qwen2-7B) and paraphrasers (Mistral-7B, Qwen2-7B, Gemma-7B), systematically stress-testing robustness under up to eight rounds of paraphrasing. Compared to prior z-score thresholding baselines and some state-of-the-art deep learning methods, our approach improves detection AUC (area under the receiver operating characteristic curve) by over 10-15 percentage points across different token lengths. Additionally, extensive cross-domain experiments demonstrate that a single universal classifier generalizes across different LLMs, paraphrasers, and text domains without retraining, maintaining above 87.8% AUC even when all components differ from training.
作者 · Sina Mansouri, Mohit Marvania, Abolfazl Safikhani