유명한 사실일수록 AI 모델에서 지우기 어렵다는 점을 반영해, 인기도에 따라 지우는 힘을 조절하는 새 방법 AdaPop
arXiv:2608.142292026-08-13
The More Popular, The Harder to Forget: Adaptive Popularity for LLM Unlearning
유명한 사실일수록 AI 모델에서 지우기 어렵다는 점을 반영해, 인기도에 따라 지우는 힘을 조절하는 새 방법 AdaPop
AI 모델에서 특정 지식을 지우는 '언러닝' 작업은 유명한 사실(사람들이 많이 아는 정보)일수록 더 깊게 학습되어 있어 지우기가 어렵다는 문제가 있다. AdaPop은 위키데이터 사이트링크 수 같은 외부 인기도 지표를 이용해 사실마다 지우는 힘을 다르게 주고, 매 에폭마다 보존해야 할 지식이 손상되지 않도록 자동으로 균형을 맞추는 조절기를 결합했다. 세 종류의 모델과 두 개의 벤치마크에서, 말을 바꿔 물어봐도 지워진 내용이 다른 방법보다 약 5배 덜 새어나오고, 일부러 캐묻는 질문에도 약 1.6배 덜 새어나왔다.
METAL MEDIA 해설 도표
AdaPop 작동 구조: 인기도 반영 지우기와 자동 균형 조절
증거 상태측정 결과가 보고됨
인기도 점수 산출위키데이터 사이트링크 수나 LLM 심사관 평가로 각 사실이 얼마나 유명한지 외부 지표에서 점수를 매긴다
지수(β)로 변환인기도 점수를 지수 β로 바꿔 유명한 사실일수록 더 세게, 드문 사실일수록 더 약하게 지우는 압력을 토큰 단위로 부여한다
가중 지우기 손실토큰 확신도와 β를 결합한 가중치를 지우기 손실에 곱해 학습 신호를 만든다
이중 상승 조절기매 에폭 끝에 보존해야 할 지식의 손실이 얼마나 늘었는지 관찰해 보존 페널티(α)를 올리거나 내린다
결과 검증말 바꿔 묻기·캐묻기 질문과 내부 은닉 상태 비교를 통해 표면적 은닉이 아닌 실제 지식 제거 여부를 확인한다
METAL MEDIA이 원문을 바탕으로 재구성한 해설 도표이며, 논문 저자의 원문 figure가 아닙니다.
무엇을 했나
문제 제기: 기존 언러닝 방법들은 지워야 할 모든 사실에 똑같은 세기로 학습을 되돌리는데, 실제로는 유명한 사실이 학습 중 더 깊게 새겨져서 잘 안 지워지고 드문 사실은 오히려 과하게 지워져 다른 지식까지 손상시킨다.
방법: AdaPop은 위키데이터 사이트링크 수나 LLM 심사관 평가 같은 외부 지표로 각 사실의 인기도를 측정한 뒤, 이를 지수(β) 형태로 바꿔 인기 있는 사실에는 더 세게, 드문 사실에는 더 약하게 지우는 압력을 준다. 동시에 매 에폭마다 보존해야 할 지식의 손실 정도를 관찰해 보존 페널티를 자동으로 조절하는 이중 상승(dual-ascent) 조절기를 붙였다.
실험: Llama-3.1-8B-Instruct, Qwen2.5-7B-Instruct, Gemma-7B-it 세 모델과 DUET, RWKU 두 벤치마크에서 GA, GD, NPO, WGA 등 기존 방법들과 비교했다.
결과: 말을 바꿔 물어보는 질문(paraphrase)에서 약 5배 덜 새어나왔고, 일부러 캐묻는 질문(adversarial)에서는 약 1.6배 덜 새어나왔다. 내부적으로도 지워진 사실의 은닉 표현이 지우기 전 모델과 더 멀어지는 반면 보존해야 할 지식의 표현은 거의 그대로 유지됐다.
Figure 1: Overview of AdaPop. The popularity score reweights per-fact token contributions to the forget loss; the dual-ascent controller adjusts the retain penalty each epoch to maintain retain quality (Algorithm 1).
Table 1: ROUGE-L Recall and Cosine Similarity for Llama-3.1-8B-Instruct at lr=10−4. ROUGE-L: mean ± std over three seeds; Cosine Similarity: single seed. F. ↓ = forget; R. ↑ = retain. w/o MU = the original checkpoint, evaluated with no unlearning applied. Underline = best among NPO, WGA, AdaPop. †Unusable: generation failure (GA: joint output and retain collapse; GD: severe retain degradation with collapse-style generation artefacts; see Appendix G).
Bench.
Algorithm
Rouge
Cos Sim
F. ↓
R. ↑
F. ↓
R. ↑
w/o MU
0.939
0.968
0.874
0.883
AdaPop
0.043 ± .007
0.959 ± .009
0.369
0.971
GA†
0.000 ± .000
0.000 ± .000
0.093
0.093
GD†
0.021 ± .003
0.853 ± .011
0.120
0.897
NPO
0.670 ± .095
0.996 ± .005
0.684
0.998
DUET
WGA
0.036 ± .002
0.995 ± .003
0.442
0.996
w/o MU
0.755
0.827
0.776
0.823
AdaPop
0.078 ± .005
0.972 ± .008
0.125
0.977
GA†
0.002 ± .000
0.001 ± .000
0.063
0.067
GD†
0.026 ± .005
0.759 ± .021
0.096
0.829
NPO
0.540 ± .034
0.957 ± .005
0.403
0.967
RWKU
WGA
0.095 ± .009
0.977 ± .003
0.247
0.984
Figure 2: Wikidata popularity-score distributions for the forget sets of DUET and RWKU. DUET spans nearly two orders of magnitude (69–3,763; median 1,090), providing a direct test of popularity-sensitive methods; RWKU is narrower and skewed lower (median 130).
Table 2: ROUGE-L Recall and Cosine Similarity for Qwen2.5-7B-Instruct at lr=10−4. Notation as in Table 1.
Bench.
Algorithm
Rouge
Cos Sim
F. ↓
R. ↑
F. ↓
R. ↑
w/o MU
0.921
0.886
0.864
0.861
AdaPop
0.056 ± .007
0.950 ± .008
0.456
0.964
GA†
0.000 ± .000
0.000 ± .000
0.058
0.092
GD†
0.022 ± .004
0.757 ± .015
0.094
0.841
NPO
0.684 ± .023
0.973 ± .002
0.576
0.979
DUET
WGA
0.058 ± .002
0.987 ± .009
0.477
0.999
w/o MU
0.570
0.666
0.322
0.347
AdaPop
0.016 ± .007
0.855 ± .003
0.178
0.911
GA†
0.001 ± .000
0.000 ± .000
0.071
0.075
GD†
0.014 ± .003
0.462 ± .018
0.150
0.640
NPO
0.271 ± .010
0.783 ± .012
0.490
0.870
RWKU
WGA
0.038 ± .016
0.890 ± .005
0.212
0.931
Figure 3: AdaPop on DUET (Llama) under three popularity signals: Wikidata score, LLM-as-Judge (3-seed mean), and measured corpus frequency (infini-gram over Pile-train). Left: forget ROUGE-L (↓). Right: holdout ROUGE-L (↑). Each point averages a rare-tier and a popular-tier run at that learning rate. All signals use the anchors of Appendix B, (a,b)=(58.7,0.796). Rates above 5×10−4 are omitted: at 10−3 every signal has at least one collapsed tier.
Table 3: ROUGE-L Recall and Cosine Similarity for Gemma-7B-it at lr=10−4. Notation as in Table 1.
Bench.
Algorithm
Rouge
Cos Sim
F. ↓
R. ↑
F. ↓
R. ↑
w/o MU
0.892
0.926
0.586
0.541
AdaPop
0.023 ± .004
0.976 ± .011
0.206
0.976
GA†
0.000 ± .000
0.000 ± .000
0.096
0.102
GD†
0.044 ± .007
0.598 ± .024
0.069
0.626
NPO
0.626 ± .071
0.968 ± .007
0.394
0.956
DUET
WGA
0.050 ± .002
0.996 ± .004
0.237
0.997
w/o MU
0.471
0.551
0.328
0.349
AdaPop
0.034 ± .014
0.948 ± .009
0.068
0.961
GA†
0.000 ± .000
0.000 ± .000
0.000
0.000
GD†
0.013 ± .003
0.334 ± .013
0.115
0.532
NPO
0.341 ± .013
0.773 ± .010
0.240
0.840
RWKU
WGA
0.040 ± .003
0.950 ± .007
0.135
0.965
Figure 4: Internal representation metrics at lr=10−4, across three models. Each row corresponds to one metric; each column corresponds to one data split. For the forget split, deeper erasure corresponds to lower ΔLP, higher ΔRank, lower Hid.Cos, and higher KL; directional desiderata are reversed for the retain split. Marker shape indicates benchmark (DUET vs. RWKU); colour indicates model family.
Table 4: Robustness at lr=10−4. Top: paraphrase ROUGE-L on the DUET forget split (lower is better). Bottom: RWKU Level-3 adversarial-attack ROUGE-L; lower indicates stronger resistance to knowledge recovery. Underline = best among NPO, WGA, AdaPop. GA and GD collapse and are excluded from forget-quality ranking.
Algo
Llama
Qwen
Gemma
DUET paraphrase forget ROUGE-L ↓
AdaPop
0.045
0.074
0.027
GA†
0.000
0.000
0.000
GD†
0.015
0.045
0.117
NPO
0.696
0.605
0.568
WGA
0.042
0.104
0.050
RWKU adversarial-attack forget ROUGE-L ↓
AdaPop
0.262
0.144
0.195
GA†
0.000
0.001
0.000
GD†
0.254
0.171
0.103
NPO
0.657
0.400
0.485
WGA
0.396
0.211
0.239
Figure 5: Component ablation on DUET. Top row: rare-fact paraphrase forget ROUGE-L (↓) and retain ROUGE-L (↑) over lr∈[2×10−5,8×10−4]. Bottom row: popular-fact paraphrase forget ROUGE-L (↓) and retain ROUGE-L (↑) over lr∈[2×10−4,8×10−3].
Table 5: DUET paraphrase forget ROUGE-L (↓) split by popularity tier at lr=10−4. Underline = best among NPO, WGA, AdaPop. GA is omitted (all values 0.000 under collapse); †GD’s low rare-tier values coincide with retain collapse (Table 1).
Algo
Llama
Qwen
Gemma
Pop.
Rare
Pop.
Rare
Pop.
Rare
AdaPop
0.040
0.049
0.055
0.092
0.028
0.026
GD†
0.027
0.003
0.086
0.003
0.232
0.002
NPO
0.854
0.538
0.778
0.432
0.796
0.339
WGA
0.067
0.017
0.194
0.014
0.096
0.003
Figure 6: ROUGE-L (solid) and Cosine Similarity (dashed) on the merged DUET forget (left) and retain (right) splits across learning rates, for Llama (top), Qwen (middle), and Gemma (bottom).
Table 6: Internal metrics at lr=10−4, averaged across Llama, Qwen, and Gemma. Directional arrows are shown in each section header. Underline = best among NPO, WGA, AdaPop. GA and GD values are reported for reference but indicate model collapse and are not ranked. Per-model breakdown in Appendix K.
Algo
ΔLP
ΔRank
Hid.Cos
KL
DUET — forget split (ΔLP↓, ΔRank↑, Hid.Cos↓, KL↑)
AdaPop
−160¯
+53,760¯
0.702
42.5
GA†
−3,106
+106,826
0.352
479.6
GD†
−5,917
+104,134
0.192
1006.4
NPO
−42
+8,824
0.861
5.4
WGA
−52
−3,714
0.729
16.2
RWKU — forget split (ΔLP↓, ΔRank↑, Hid.Cos↓, KL↑)
AdaPop
−23¯
+4,169¯
0.484
9.1
GA†
−6,083
+96,254
0.285
953.1
GD†
−6,842
+96,881
0.175
1139.7
NPO
−18
+2,739
0.883
2.9
WGA
−3.5
−11,010
0.541
6.7
Retain split — DUET (ΔLP↑, ΔRank↓, Hid.Cos↑, KL↓)
AdaPop
+21
−13,770
0.875
6.5
GA†
−2,750
+103,104
0.348
511.4
GD†
−251
−8,751
0.760
48.2
NPO
+21
−13,660
0.910
7.5
WGA
+23¯
−13,930¯
0.880
7.5
Retain split — RWKU (ΔLP↑, ΔRank↓, Hid.Cos↑, KL↓)
AdaPop
+33
−12,190
0.883
6.2
GA†
−6,106
+98,802
0.289
955.8
GD†
−540
−3,122
0.783
97.1
NPO
+31
−11,727
0.869
7.4
WGA
+34¯
−12,202¯
0.875
6.8
Figure 7: ROUGE-L (solid) and Cosine Similarity (dashed) on the RWKU forget (left) and retain (right) splits across learning rates, for Llama (top), Qwen (middle), and Gemma (bottom).
Table 7: MMLU accuracy and HellaSwag (HS) normalised accuracy at lr=10−4. w/o MU = the original checkpoint, with no unlearning applied. Values averaged across DUET and RWKU checkpoints.
Algo
Llama
Qwen
Gemma
MMLU
HS
MMLU
HS
MMLU
HS
w/o MU
0.65
0.73
0.71
0.68
0.47
0.64
AdaPop
0.65
0.76
0.71
0.68
0.52
0.62
GA
0.23
0.33
0.23
0.43
0.25
0.26
GD
0.64
0.59
0.71
0.65
0.47
0.37
NPO
0.65
0.72
0.70
0.68
0.51
0.63
WGA
0.66
0.76
0.71
0.70
0.52
0.68
Figure 8: DUET forget and retain splits separated by popularity tier (Llama). Top row: rare facts. Bottom row: popular facts. The order-of-magnitude difference in the learning rate required to achieve comparable forgetting across the two tiers is direct evidence of the popularity gap.
Table 8: AdaPop coefficient sensitivity. Each row perturbs a and/or b by ±20% from the analytically derived baseline. F. ↓ = forget ROUGE-L; R. ↑ = retain ROUGE-L. Δ values are signed differences from the baseline.
Config
a
b
F. ↓
R. ↑
ΔF.
ΔR.
baseline
58.70
0.7960
0.045
0.961
0
0
+a,+b
70.44
0.9552
0.038
0.975
−0.007
+0.014
+a,−b
70.44
0.6368
0.044
0.994
−0.001
+0.033
−a,+b
46.96
0.9552
0.024
0.920
−0.021
−0.041
−a,−b
46.96
0.6368
0.034
0.991
−0.011
+0.030
Figure 9: Forget (left) and retain (right) ROUGE-L Recall per training epoch on DUET, Llama-3.1-8B-Instruct, at lr=10−4.
Table 9: Discordant cases between Wikidata score and LLM-judge scores on DUET. ROUGE-L: Llama-3.1-8B-Instruct recall before unlearning. Top group: Wikidata-popular facts recalled at ROUGE-L 1.0 despite low LLM-judged salience, indicating deep memorisation. Bottom group: Wikidata-rare facts with higher LLM scores but variable recall, consistent with shallower encoding.
Question
Answer
Wikidata
LLM
ROUGE-L
Wikidata popular, LLM rare
What is the country of Ensenada?
Mexico
2113
30
1.0
What is the country of Tourcoing?
France
2112
40
1.0
What is the country of Qus?
Egypt
3763
50
1.0
Wikidata rare, LLM popular
What is the instance of Hua Hin?
seaside resort
148
9000
0.5
What is the instance of Weitra?
municipality of Austria
141
5000
0.7
What is the located in the administrative territorial entity of Ollantaytambo?
Urubamba Province
165
4000
1.0
Table 10: Agreement of each popularity proxy with measured corpus frequency on DUET, counted over Pile-train (383B tokens) with infini-gram. Pearson is computed in log space. The two proxies agree with each other at Spearman 0.596 and assign the same rare/popular label to 84% of facts.
Proxy
Pearson
Spearman
Wikidata score
0.970
0.839
LLM judge
0.677
0.659
Table 11: AdaPop under corrupted popularity scores (DUET, Llama-3.1-8B-Instruct, lr=10−4). Forget ROUGE-L is reported per popularity tier. Retain degrades by 0.04 even when every label is inverted.
Score noise
F. rare ↓
F. pop. ↓
R. ↑
none
0.05
0.04
0.96
log-normal, σ=1.0
0.02
0.05
0.96
labels 50% swapped
0.02
0.05
0.97
labels 100% swapped
0.01
0.05
0.92
Table 12: ROUGE-L Recall for additional baselines, Llama-3.1-8B-Instruct at lr=10−4. w/o MU = the original checkpoint, with no unlearning applied; AdaPop reproduced from Table 1 for reference. F. ↓ = forget; R. ↑ = retain. Underline = best among non-collapsed methods. †Model collapse.
Bench.
Algorithm
ROUGE-L
F. ↓
R. ↑
w/o MU
0.939
0.968
AdaPop
0.043
0.959
UNDIAL
0.884
0.998
RMU
0.882
0.998
PDU
0.476
0.954
NPO-SAM
0.684
0.970
SimNPO
0.340
0.999
SatImp
0.108
0.995
Adaptive RMU
0.933
0.963
AltPO
0.181
1.000
FLAT†
0.001
0.001
TPO
0.397
0.997
DUET
CE-U†
0.000
0.000
w/o MU
0.755
0.827
AdaPop
0.078
0.972
UNDIAL
0.568
0.939
RMU
0.801
0.986
PDU
0.201
0.909
NPO-SAM
0.529
0.885
SimNPO
0.573
0.987
SatImp
0.417
0.985
Adaptive RMU
0.162
0.819
AltPO
0.200
0.988
FLAT†
0.004
0.003
TPO
0.341
0.986
RWKU
CE-U†
0.000
0.000
Table 13: ROUGE-L Recall for additional baselines, Qwen2.5-7B-Instruct at lr=10−4. Notation as in Table 12.
Bench.
Algorithm
ROUGE-L
F. ↓
R. ↑
w/o MU
0.921
0.886
AdaPop
0.056
0.950
UNDIAL
0.797
0.930
RMU
0.734
0.983
PDU
0.632
0.869
NPO-SAM
0.211
0.769
SimNPO
0.295
0.991
SatImp
0.070
0.985
Adaptive RMU
0.562
0.831
AltPO
0.247
0.999
FLAT†
0.000
0.001
TPO
0.460
0.999
DUET
CE-U†
0.001
0.001
w/o MU
0.570
0.666
AdaPop
0.016
0.855
UNDIAL
0.313
0.668
RMU
0.407
0.922
PDU
0.455
0.710
NPO-SAM
0.298
0.596
SimNPO
0.338
0.934
SatImp
0.187
0.891
Adaptive RMU
0.118
0.466
AltPO
0.170
0.989
FLAT†
0.004
0.003
TPO
0.382
0.934
RWKU
CE-U†
0.003
0.001
Table 14: ROUGE-L Recall for additional baselines, Gemma-7B-it at lr=10−4. Notation as in Table 12.
Bench.
Algorithm
ROUGE-L
F. ↓
R. ↑
w/o MU
0.892
0.926
AdaPop
0.023
0.976
UNDIAL
0.819
0.980
RMU
0.851
0.992
PDU
0.677
0.934
NPO-SAM
0.539
0.777
SimNPO
0.581
0.997
SatImp
0.070
0.998
Adaptive RMU
0.293
0.828
AltPO
0.262
1.000
FLAT†
0.001
0.001
TPO
0.899
0.999
DUET
CE-U†
0.000
0.000
w/o MU
0.471
0.551
AdaPop
0.034
0.948
UNDIAL
0.453
0.830
RMU
0.547
0.968
PDU
0.404
0.825
NPO-SAM
0.271
0.644
SimNPO
0.373
0.969
SatImp
0.102
0.964
Adaptive RMU
0.091
0.535
AltPO
0.172
0.989
FLAT†
0.005
0.004
TPO
0.594
0.971
RWKU
CE-U†
0.000
0.000
실제로 확인된 결과
말을 바꿔 물어보는 질문(paraphrase)에서 경쟁 방법들 대비 약 5배 덜 새어나왔고, 일부러 캐묻는 질문(adversarial reformulation)에서는 약 1.6배 덜 새어나왔다.
세 모델(Llama, Qwen, Gemma)과 두 벤치마크(DUET, RWKU)에서 학습률 10^-4 기준, 안정적으로 작동하는 방법들 중 지워야 할 항목의 코사인 유사도가 6개 조합 모두에서 가장 낮았고, ROUGE-L도 6개 중 5개에서 가장 낮았다.
일반 능력을 측정하는 MMLU 점수는 지우기 전 모델과 0.05 이내 차이를 유지했다.
위키데이터 점수, LLM 심사관 점수, 실제 말뭉치 빈도(Pile-train, infini-gram) 세 가지 인기도 지표는 서로 84%의 사실에서 같은 희귀/인기 분류에 합의했고, 인기도 점수를 완전히 뒤집어도(모든 라벨 역전) 보존 성능은 0.04만 떨어졌다.
어디에 쓸 수 있나
개인정보나 저작권 침해 등의 이유로 특정 사실을 언어모델에서 제거해야 하는 서비스 운영 상황
위키데이터처럼 이미 존재하는 대중적 인지도 지표를 활용해 지식 삭제 강도를 자동 조절하려는 연구·개발
말을 바꿔 묻거나 캐묻는 질문에도 지워진 정보가 다시 드러나지 않도록 검증하는 언러닝 평가 파이프라인 설계
한계와 남은 검증
실험은 7~8B 규모의 세 모델(Llama-3.1-8B-Instruct, Qwen2.5-7B-Instruct, Gemma-7B-it)과 두 벤치마크(DUET, RWKU)에 한정되어 검증되었고 더 크거나 다른 구조의 모델에서는 확인되지 않았다.
LoRA 파인튜닝만 사용했고 전체 파인튜닝은 다른 선행 연구의 결과를 근거로 제외되어 별도로 검증되지 않았다.
인기도 지표로 주로 위키데이터 점수를 썼고, 위키데이터 커버리지가 없는 사실에 대한 LLM 심사관 방식의 안정성은 상대적으로 약하게만 확인되었다.
Llama/DUET 조합에서는 WGA라는 다른 방법이 표면적인 ROUGE-L 지표에서 근소하게 더 낮은 값을 보이는 예외가 있었다.
왜 중요한가
프라이버시나 안전, 법적 준수를 위해 AI 모델에서 특정 정보를 지워야 하는 상황이 늘고 있는데, 단순히 답변만 숨기는 것이 아니라 실제로 모델 내부에서 그 지식을 없애는지가 중요하다. 이 연구는 '얼마나 유명한 사실인가'라는 정보를 활용하면 더 확실하게, 그리고 다른 지식을 덜 손상시키며 지울 수 있다는 것을 보여준다.
이 논문의 용어
언러닝(unlearning) · 학습을 처음부터 다시 하지 않고 이미 학습된 모델에서 특정 지식만 제거하는 작업
위키데이터 사이트링크 점수 · 위키데이터의 한 항목이 몇 개 언어판 위키백과에 연결되어 있는지를 세어 유명세를 가늠하는 지표
이중 상승(dual-ascent) 조절기 · 지우는 힘과 보존 페널티 사이 균형을 매 에폭마다 관찰된 손상 정도에 따라 자동으로 조정하는 장치
ROUGE-L · 모델이 생성한 답과 정답 사이 겹치는 단어 순서를 재는 지표. 지우려는 지식에서는 낮을수록, 보존할 지식에서는 높을수록 좋다
LLM-as-Judge · 큰 언어모델에게 특정 사실이 얼마나 유명한지 평가하도록 시켜 얻는 점수