Stability-Aware Feature Design for Robust Watermark Detection in Machine-Generated Text
AI가 쓴 티를 내는 워터마크, 문장을 여러 번 바꿔써도 잡아내는 탐지법
AI가 생성한 텍스트에는 눈에 안 띄게 심어둔 통계적 신호(워터마크)가 있는데, 사람이 다른 표현으로 여러 번 바꿔쓰면(패러프레이징) 기존 탐지기는 성능이 확 떨어진다. 이 논문은 텍스트를 여러 구간으로 쪼개 국소적으로 신호를 살피고, 문장을 반복해서 바꿔써도 그 신호가 얼마나 일정하게 유지되는지를 측정하는 PSS(Pattern Stability Score)라는 새 특징을 만들었다. 그 결과 8번 바꿔쓴 극단적인 상황에서도 기존 방식보다 훨씬 안정적으로 AI 생성 여부를 구분해냈다.
METAL MEDIA 해설 도표
AI가 쓴 티를 내는 워터마크, 문장을 여러 번 바꿔써도 잡아내는 탐지법
01AI 생성 텍스트를 표시하는 워터마크는 단어 선택을 '초록 리스트'로 살짝 편향시키는 방식(greenlist watermarking)인데, 기존 탐지기는 전체 텍스트를 하나의 점수(z-score)로 뭉뚱그려 판단해서 일부 구간만 편집돼도 전체 점수가 희석되는 문제가 있었다
02연구팀은 텍스트를 작은 창(window)으로 나눠 국소적으로 통계치(평균, 분산, 자기상관, 연속 구간 길이 등 20가지 특징)를 뽑고, 같은 텍스트를 최대 8번 반복해서 바꿔쓴 뒤에도 이 국소 신호가 얼마나 안정적으로 유지되는지를 계산하는 PSS를 도입했다
03PG-19, CNN/DailyMail, WikiText 세 데이터셋과 Llama-3-8B, Qwen2-7B 등 여러 생성 모델, Mistral·Qwen2·Gemma 등 여러 패러프레이저로 실험한 결과, 1500토큰 기준 8번 바꿔쓴 텍스트에서도 PSS 방식은 91.2%의 판별 정확도(AUC)를 유지한 반면 기존 딥러닝 기반 탐지기(DeepTextMark, Binoculars, RADAR)는 41~44%로 무너졌다
04짧은 글(300~500토큰)이나 서로 다른 데이터셋·모델 조합에서도, 재학습 없이 하나의 분류기로 87.8% 이상의 AUC를 유지하는 범용성도 확인했다
05다만 텍스트가 아주 짧으면서 동시에 깊게 반복해서 바꿔쓴 경우, 그리고 워터마크 신호 자체를 없애도록 특별히 학습된 공격 앞에서는 성능이 크게 떨어지는 한계도 있었다
METAL MEDIA이 원문을 바탕으로 재구성한 해설 도표이며, 논문 저자의 원문 figure가 아닙니다.
무엇을 했나
AI 생성 텍스트를 표시하는 워터마크는 단어 선택을 '초록 리스트'로 살짝 편향시키는 방식(greenlist watermarking)인데, 기존 탐지기는 전체 텍스트를 하나의 점수(z-score)로 뭉뚱그려 판단해서 일부 구간만 편집돼도 전체 점수가 희석되는 문제가 있었다
연구팀은 텍스트를 작은 창(window)으로 나눠 국소적으로 통계치(평균, 분산, 자기상관, 연속 구간 길이 등 20가지 특징)를 뽑고, 같은 텍스트를 최대 8번 반복해서 바꿔쓴 뒤에도 이 국소 신호가 얼마나 안정적으로 유지되는지를 계산하는 PSS를 도입했다
PG-19, CNN/DailyMail, WikiText 세 데이터셋과 Llama-3-8B, Qwen2-7B 등 여러 생성 모델, Mistral·Qwen2·Gemma 등 여러 패러프레이저로 실험한 결과, 1500토큰 기준 8번 바꿔쓴 텍스트에서도 PSS 방식은 91.2%의 판별 정확도(AUC)를 유지한 반면 기존 딥러닝 기반 탐지기(DeepTextMark, Binoculars, RADAR)는 41~44%로 무너졌다
짧은 글(300~500토큰)이나 서로 다른 데이터셋·모델 조합에서도, 재학습 없이 하나의 분류기로 87.8% 이상의 AUC를 유지하는 범용성도 확인했다
다만 텍스트가 아주 짧으면서 동시에 깊게 반복해서 바꿔쓴 경우, 그리고 워터마크 신호 자체를 없애도록 특별히 학습된 공격 앞에서는 성능이 크게 떨어지는 한계도 있었다
Figure 2: AUC vs. paraphrase depth at 1,500 tokens. Mean AUC (solid lines) with ±1 SD bands (shaded) over 30 runs with random 70/30 splits for four methods—Global z threshold, Local z (20) + classifiers, Static features + classifiers, and PSS + Static—all using XGBoost. Local statistics improve robustness relative to the global threshold, and adding PSS further flattens the AUC decline from D3–D8.
Table 1: Comparison with prior detectors on PG-19 at 1,500 tokens. AUC (%) across paraphrase depths. PSS + Static outperforms statistical baselines (Global z-score, WinMax) by 14–19 percentage points an, the global z-score by 22–24 percentage points, and deep-learning detectors (DeepTextMark, Binoculars, RADAR) by 30–50 percentage points across D1–D8.
Method
D1
D3
D5
D8
Global z-score
74.2
70.0
68.0
66.8
WinMax (Kirchenbauer et al., 2024)
82.1
76.5
74.4
72.3
Static features (20-D)
88.1
83.2
80.2
78.6
DeepTextMark (Munyer et al., 2024)
65.8
51.2
47.5
43.9
Binoculars (Hans et al., 2024)
62.7
48.6
44.9
41.9
RADAR (Hu et al., 2023)
58.4
47.2
43.7
40.8
PSS + Static
96.1
93.9
92.6
91.2
Figure 3: Window/stride sensitivity (1,500 tokens) for PSS + static. Numbers show mean AUC (%) across depths D1–D8.
Table 2: Cross-Domain Transfer. AUC (%) when training on one dataset and testing on another. PSS generalizes effectively while deep learning methods collapse.
Train → Test
Method
D1
D3
D5
D8
PG-19 → CNN/DM
DeepTextMark
46.3
38.5
34.2
32.1
Binoculars
44.1
36.8
33.1
30.9
RADAR
41.8
35.1
31.5
29.4
PSS + Static
88.6
85.8
84.7
83.8
CNN/DM → WikiText
DeepTextMark
50.8
41.2
37.5
35.2
Binoculars
48.9
39.8
36.2
34.0
RADAR
46.7
38.2
34.8
32.6
PSS + Static
93.1
91.2
90.4
89.7
WikiText → CNN/DM
DeepTextMark
43.2
35.4
31.3
28.9
Binoculars
41.5
33.9
29.8
27.6
RADAR
39.4
32.1
28.1
26.2
PSS + Static
88.6
85.7
84.6
83.7
Figure 4: Shorter texts. Mean AUC (solid) with ±1 SD bands (shaded) over 30 random 70/30 splits. Top-left to bottom-left: AUC vs. paraphrase depth for 1,000/500/300 tokens; bottom-right: AUC vs. token length at D7. All methods degrade with less text, but local/static features mitigate the drop and PSS + Static maintains the strongest performance across depths and lengths, including at D7.
Table 3: Universal Classifier Performance. AUC (%) for a single classifier trained on Llama-3 + Mistral + PG-19 + D1-D3, evaluated across varied configurations.
Test Configuration
Changed
D1
D3
D5
D8
Llama-3 + Mistral + PG-19
Baseline
96.1
93.9
92.6
91.2
Qwen2 + Mistral + PG-19
LLM
94.2
91.8
90.9
89.6
Llama-3 + Gemma + PG-19
Paraphraser
89.1
82.7
80.6
80.0
Llama-3 + Mistral + CNN
Domain
88.6
85.8
84.7
83.8
Qwen2 + Qwen2 + WikiText
ALL
92.6
90.5
89.2
87.8
Figure 5: Mix paraphrasing (Mistral ↔ Qwen), 1,500 tokens. AUC vs. depth under alternating paraphrasers.
Table 4: Gemma-7B-IT paraphrasing: AUC (%) vs. depth (D1–D8). All classifier entries use XGBoost; values are mean ± std over 30 runs.
Method
D1
D2
D3
D4
D5
D6
D7
D8
Global z-score threshold
56.50 ± 0.00
54.35 ± 0.00
53.75 ± 0.00
53.55 ± 0.00
53.25 ± 0.00
53.30 ± 0.00
53.05 ± 0.00
53.00 ± 0.00
Local z-score (20)
66.07 ± 1.42
59.03 ± 1.66
58.92 ± 1.99
57.52 ± 2.20
57.81 ± 1.63
57.32 ± 1.40
57.55 ± 1.55
57.80 ± 1.41
Static features
72.75 ± 1.87
68.95 ± 1.87
67.02 ± 1.29
68.10 ± 1.35
67.48 ± 1.76
67.10 ± 1.83
66.82 ± 1.58
65.16 ± 1.87
PSS + Static
83.51 ± 1.43
77.85 ± 1.63
75.98 ± 1.45
76.54 ± 1.21
73.69 ± 1.49
73.33 ± 1.23
72.92 ± 2.03
–
Table 5: Qwen2-7B-Instruct paraphrasing: AUC (%) vs. depth (D1–D8). All classifier entries use XGBoost; values are mean ± std over 30 runs.
Method
D1
D2
D3
D4
D5
D6
D7
D8
Global z-score threshold
73.25 ± 0.00
69.90 ± 0.00
69.00 ± 0.00
68.40 ± 0.00
67.90 ± 0.00
67.70 ± 0.00
67.35 ± 0.00
66.80 ± 0.00
Local z-score (20)
58.90 ± 1.61
58.15 ± 1.69
55.03 ± 1.71
55.48 ± 1.64
55.86 ± 1.66
55.19 ± 1.64
54.27 ± 1.63
55.27 ± 1.66
Static features
82.08 ± 1.45
78.64 ± 1.55
77.15 ± 1.51
77.89 ± 1.48
77.47 ± 1.52
78.37 ± 1.50
78.77 ± 1.49
76.36 ± 1.47
PSS + Static
94.67 ± 0.68
92.68 ± 0.79
92.38 ± 0.82
92.63 ± 0.78
91.90 ± 0.91
90.57 ± 0.98
90.47 ± 1.02
–
Table 6: Feature-group ablation. AUC (%) on PG-19 at 1,500 tokens. Components are added incrementally to a global z-score baseline. Every group contributes measurably and PSS provides the largest single jump, confirming that each component captures a distinct, non-redundant aspect of the watermark signal.
Feature Set
D1
D3
D5
D8
Global z-score (baseline)
74.2
70.0
68.0
66.8
Z-score moments (6-D)
83.2
76.5
74.1
72.0
+ Autocorrelations (8-D)
85.0
79.3
76.0
74.4
+ Run-length stats (14-D)
87.3
83.2
79.8
77.2
+ Run-frequency (full 20-D)
88.1
83.2
80.2
78.6
+ PSS (full PSS + Static)
96.1
93.9
92.6
91.2
Table 7: PSS with Different Watermarking Schemes. AUC (%) with varying greenlist ratios.
Watermark Config
D1
D3
D5
D8
Standard (γ=0.25)
96.1
93.9
92.6
91.2
Stronger (γ=0.50)
97.6
96.4
95.5
94.2
Table 8: Within-domain: Train on CNN/DailyMail (70%) → Test on CNN/DailyMail (30%). AUC (%) across paraphrase depths. All classifier entries use XGBoost; values are mean ± std over 30 runs.
Method
D1
D2
D3
D4
D5
D6
D7
D8
Global z-score threshold
74.20 ± 0.00
70.45 ± 0.00
69.60 ± 0.00
68.85 ± 0.00
68.20 ± 0.00
67.70 ± 0.00
67.30 ± 0.00
66.80 ± 0.00
Local z-score (20)
59.80 ± 1.78
58.65 ± 1.82
56.40 ± 1.75
55.50 ± 1.77
54.90 ± 1.80
54.35 ± 1.83
53.80 ± 1.85
53.25 ± 1.88
Static features
84.30 ± 1.58
82.50 ± 1.62
78.05 ± 1.60
77.65 ± 1.65
77.05 ± 1.68
76.45 ± 1.72
75.85 ± 1.75
75.25 ± 1.78
PSS + Static
95.30 ± 0.85
93.50 ± 0.92
92.75 ± 0.98
92.35 ± 1.05
91.70 ± 1.12
91.15 ± 1.18
90.60 ± 1.25
90.05 ± 1.32
Table 9: Within-domain: Train on WikiText (70%) → Test on WikiText (30%). AUC (%) across paraphrase depths. All classifier entries use XGBoost; values are mean ± std over 30 runs.
Method
D1
D2
D3
D4
D5
D6
D7
D8
Global z-score threshold
79.85 ± 0.00
77.15 ± 0.00
76.45 ± 0.00
75.90 ± 0.00
75.45 ± 0.00
75.10 ± 0.00
74.80 ± 0.00
74.55 ± 0.00
Local z-score (20)
65.70 ± 1.55
64.90 ± 1.58
63.40 ± 1.62
62.85 ± 1.60
62.50 ± 1.63
62.20 ± 1.65
61.90 ± 1.67
61.65 ± 1.68
Static features
88.85 ± 1.48
88.10 ± 1.50
85.80 ± 1.46
85.45 ± 1.48
85.15 ± 1.50
84.85 ± 1.52
84.55 ± 1.54
84.30 ± 1.55
PSS + Static
98.45 ± 0.55
97.25 ± 0.62
96.85 ± 0.68
96.50 ± 0.72
96.15 ± 0.78
95.85 ± 0.82
95.55 ± 0.85
95.30 ± 0.88
Table 10: Cross-domain: Train on WikiText (70%) → Test on CNN/DailyMail (30%). AUC (%) across paraphrase depths. All classifier entries use XGBoost; values are mean ± std over 30 runs.
Method
D1
D2
D3
D4
D5
D6
D7
D8
Global z-score threshold
59.80 ± 0.00
57.20 ± 0.00
55.95 ± 0.00
55.05 ± 0.00
54.40 ± 0.00
53.95 ± 0.00
53.60 ± 0.00
53.30 ± 0.00
Local z-score (20)
54.85 ± 2.02
53.55 ± 2.10
52.10 ± 2.18
51.30 ± 2.24
50.75 ± 2.28
50.30 ± 2.32
49.95 ± 2.35
49.65 ± 2.38
Static features
76.15 ± 1.82
74.20 ± 1.90
72.75 ± 1.98
72.00 ± 2.05
71.40 ± 2.10
70.95 ± 2.15
70.55 ± 2.18
70.25 ± 2.20
PSS + Static
88.60 ± 1.12
87.15 ± 1.22
85.70 ± 1.30
85.15 ± 1.38
84.60 ± 1.45
84.15 ± 1.52
83.90 ± 1.58
83.70 ± 1.65
Table 11: Cross-domain: Train on CNN/DailyMail (70%) → Test on WikiText (30%). AUC (%) across paraphrase depths. All classifier entries use XGBoost; values are mean ± std over 30 runs.
Method
D1
D2
D3
D4
D5
D6
D7
D8
Global z-score threshold
65.55 ± 0.00
63.40 ± 0.00
62.40 ± 0.00
61.70 ± 0.00
61.20 ± 0.00
60.80 ± 0.00
60.50 ± 0.00
60.25 ± 0.00
Local z-score (20)
59.35 ± 1.72
58.15 ± 1.78
56.80 ± 1.85
56.10 ± 1.90
55.55 ± 1.94
55.10 ± 1.97
54.70 ± 2.00
54.40 ± 2.02
Static features
81.45 ± 1.55
79.80 ± 1.62
78.50 ± 1.68
77.90 ± 1.72
77.40 ± 1.75
77.00 ± 1.78
76.65 ± 1.80
76.40 ± 1.82
PSS + Static
93.10 ± 0.85
92.15 ± 0.92
91.20 ± 0.98
90.80 ± 1.05
90.40 ± 1.12
90.10 ± 1.18
89.85 ± 1.25
89.70 ± 1.30
Table 12: Comprehensive Experimental Coverage. Performance across different datasets, LLMs, paraphrasers, and watermark settings.
Dataset
LLM
Paraphraser
Setting
D1
D3
D5
D8
PG-19
Llama-3-8B
Mistral-7B
γ=0.25
96.1%
93.9%
92.6%
91.2%
PG-19
Qwen2-7B
Mistral-7B
γ=0.25
94.2%
91.8%
90.9%
89.6%
CNN/DailyMail
Llama-3-8B
Mistral-7B
γ=0.25
95.3%
93.2%
92.3%
90.7%
WikiText
Llama-3-8B
Qwen2-7B
γ=0.25
97.5%
96.2%
95.8%
94.8%
PG-19
Llama-3-8B
Mistral-7B
γ=0.50
97.6%
96.4%
95.5%
94.2%
Table 13: True Positive Rate at Fixed False Positive Rates. Critical operating points for practical deployment.
FPR
Method
D1
D2
D3
D4
D5
D6
D7
D8
1%
Global z-score
0.42
0.38
0.35
0.32
0.30
0.28
0.26
0.24
Local z-score (20)
0.48
0.44
0.40
0.37
0.35
0.33
0.31
0.29
Static features
0.65
0.60
0.55
0.52
0.49
0.47
0.45
0.43
PSS + Static
0.84
0.80
0.76
0.73
0.70
0.68
0.66
0.64
5%
Global z-score
0.58
0.52
0.47
0.43
0.40
0.37
0.35
0.33
Local z-score (20)
0.62
0.57
0.52
0.48
0.45
0.42
0.40
0.38
Static features
0.78
0.73
0.68
0.64
0.61
0.58
0.56
0.54
PSS + Static
0.92
0.89
0.86
0.83
0.81
0.79
0.77
0.75
Table 14: Semantic Preservation and Performance. BERT similarity decreases with depth; PSS remains robust across all depths. All detection values are AUC (%).
Depth
BERT Sim.
PSS+Static
Global z
DeepTextMark
D1
0.92
96.1%
74.2%
58.4%
D2
0.89
94.8%
71.5%
53.1%
D3
0.85
93.9%
68.7%
47.6%
D4
0.82
93.2%
67.8%
45.8%
D5
0.80
92.6%
66.8%
43.9%
D6
0.76
92.1%
66.2%
42.8%
D7
0.72
91.6%
65.8%
42.0%
D8
0.68
91.2%
65.3%
41.2%
Table 15: Performance Under Realistic Attacks. AUC (%) across various attack scenarios. PSS maintains robust performance.
Attack Type
PSS+Static
Global z
DeepTextMark
1-2 paraphrases (D1-D2 avg)
95.4%
72.8%
54.6%
Manual edits (20% modified)
91.8%
62.7%
46.3%
Manual edits (30% modified)
88.9%
56.4%
40.2%
Chain paraphrasing (Mixed)
90.6%
67.1%
49.1%
Table 16: Computation Time Breakdown (seconds). Mean ± std over 100 runs. PSS detection is highly efficient.
Component
300 tok
500 tok
1000 tok
1500 tok
Binary sequence extraction
0.12±0.01
0.18±0.01
0.28±0.02
0.41±0.02
Rolling-window features
0.35±0.03
0.58±0.04
0.95±0.06
1.45±0.08
Stability score computation
0.28±0.02
0.46±0.03
0.78±0.05
1.18±0.07
XGBoost classification
0.05±0.01
0.08±0.01
0.09±0.01
0.16±0.02
Feature extraction and classification
0.80±0.04
1.30±0.05
2.10±0.08
3.20±0.11
Table 17: Deployment Trade-offs. Practical configurations balancing accuracy and latency.
Configuration
Paraphrases
Inference Time
D8 AUC
Use Case
Static features only
0
0.8–3.2s
83–85%
High-throughput screening
PSS (minimal)
1
20–35s
89–91%
Balanced deployment
PSS (full)
7
2–4 min
91–96%
High-stakes verification
Baselines
Global z-score
0
<0.1s
62–66%
–
DeepTextMark
0
2–5s
41–44%
–
Table 18: Performance under the adaptive attack of Diaa et al. (2025). AUC (%) on PG-19 at 1,500 tokens, comparing naive Mistral paraphrasing at D1 to the DPO-optimized adaptive attack at D1 and D2. The adaptive paraphraser is fine-tuned to minimize the global z-score and effectively defeats it. In contrast, PSS + Static maintains >80% AUC, retaining a 30+ percentage point advantage.
Naive
Adaptive (Diaa et al.)
Method
D1
D1
D2
Global z-score
74.2
52.1
53.6
PSS + Static
96.1
83.2
80.6
왜 중요한가
AI가 만든 글을 표시해두는 워터마크 기술이 실제로 쓰이려면 사람이 문장을 살짝 다듬거나 여러 번 바꿔써도 흔적이 남아야 하는데, 이 연구는 탐지기 쪽만 손봐서 그 문제를 해결하는 실용적인 방법을 보여준다. 교육, 언론, 저작권 검증처럼 AI 생성물 여부를 가려야 하는 현장에서 바로 적용 가능한 접근이라는 점에서 의미가 있다.
이 논문의 용어
워터마킹(watermarking) · AI가 생성한 문장에 사람 눈에는 안 보이지만 통계적으로 검출 가능한 표식을 몰래 심어두는 기법
greenlist 워터마킹 · 단어를 생성할 때마다 어휘 일부를 '초록 리스트'로 정해 그쪽 단어가 더 자주 나오게 살짝 편향시키는 방식
z-score · 관측된 값이 우연히 나올 확률적 기준에서 얼마나 벗어났는지를 나타내는 통계 수치, 워터마크 탐지의 기본 판별 기준
패러프레이징(paraphrasing) · 같은 의미를 다른 표현으로 바꿔 쓰는 것, 워터마크를 지우려는 대표적 공격 방법
AUC(area under the curve) · 탐지 모델이 진짜와 가짜를 얼마나 잘 구분하는지 나타내는 성능 지표, 1에 가까울수록 정확
본문에 싣지 못한 그림
Figure 1: End-to-end PSS pipeline. Human and AI-watermarked texts undergo up to eight paraphrasing rounds, with binary sequences extracted at each iteration (D0–D8). Watermarked text maintains stable patterns across iterations while human text shows random variation. An XGBoost classifier uses these stability features alongside static features for final classification.