Figure 1: Overview of AdaPop. The popularity score reweights per-fact token contributions to the forget loss; the dual-ascent controller adjusts the retain penalty each epoch to maintain retain quality (Algorithm 1).
Table 1: ROUGE-L Recall and Cosine Similarity for Llama-3.1-8B-Instruct at lr=10−4. ROUGE-L: mean ± std over three seeds; Cosine Similarity: single seed. F. ↓ = forget; R. ↑ = retain. w/o MU = the original checkpoint, evaluated with no unlearning applied. Underline = best among NPO, WGA, AdaPop. †Unusable: generation failure (GA: joint output and retain collapse; GD: severe retain degradation with collapse-style generation artefacts; see Appendix G).
Bench.
Algorithm
Rouge
Cos Sim
F. ↓
R. ↑
F. ↓
R. ↑
w/o MU
0.939
0.968
0.874
0.883
AdaPop
0.043 ± .007
0.959 ± .009
0.369
0.971
GA†
0.000 ± .000
0.000 ± .000
0.093
0.093
GD†
0.021 ± .003
0.853 ± .011
0.120
0.897
NPO
0.670 ± .095
0.996 ± .005
0.684
0.998
DUET
WGA
0.036 ± .002
0.995 ± .003
0.442
0.996
w/o MU
0.755
0.827
0.776
0.823
AdaPop
0.078 ± .005
0.972 ± .008
0.125
0.977
GA†
0.002 ± .000
0.001 ± .000
0.063
0.067
GD†
0.026 ± .005
0.759 ± .021
0.096
0.829
NPO
0.540 ± .034
0.957 ± .005
0.403
0.967
RWKU
WGA
0.095 ± .009
0.977 ± .003
0.247
0.984
Figure 2: Wikidata popularity-score distributions for the forget sets of DUET and RWKU. DUET spans nearly two orders of magnitude (69–3,763; median 1,090), providing a direct test of popularity-sensitive methods; RWKU is narrower and skewed lower (median 130).
Table 2: ROUGE-L Recall and Cosine Similarity for Qwen2.5-7B-Instruct at lr=10−4. Notation as in Table 1.
Bench.
Algorithm
Rouge
Cos Sim
F. ↓
R. ↑
F. ↓
R. ↑
w/o MU
0.921
0.886
0.864
0.861
AdaPop
0.056 ± .007
0.950 ± .008
0.456
0.964
GA†
0.000 ± .000
0.000 ± .000
0.058
0.092
GD†
0.022 ± .004
0.757 ± .015
0.094
0.841
NPO
0.684 ± .023
0.973 ± .002
0.576
0.979
DUET
WGA
0.058 ± .002
0.987 ± .009
0.477
0.999
w/o MU
0.570
0.666
0.322
0.347
AdaPop
0.016 ± .007
0.855 ± .003
0.178
0.911
GA†
0.001 ± .000
0.000 ± .000
0.071
0.075
GD†
0.014 ± .003
0.462 ± .018
0.150
0.640
NPO
0.271 ± .010
0.783 ± .012
0.490
0.870
RWKU
WGA
0.038 ± .016
0.890 ± .005
0.212
0.931
Figure 3: AdaPop on DUET (Llama) under three popularity signals: Wikidata score, LLM-as-Judge (3-seed mean), and measured corpus frequency (infini-gram over Pile-train). Left: forget ROUGE-L (↓). Right: holdout ROUGE-L (↑). Each point averages a rare-tier and a popular-tier run at that learning rate. All signals use the anchors of Appendix B, (a,b)=(58.7,0.796). Rates above 5×10−4 are omitted: at 10−3 every signal has at least one collapsed tier.
Table 3: ROUGE-L Recall and Cosine Similarity for Gemma-7B-it at lr=10−4. Notation as in Table 1.
Bench.
Algorithm
Rouge
Cos Sim
F. ↓
R. ↑
F. ↓
R. ↑
w/o MU
0.892
0.926
0.586
0.541
AdaPop
0.023 ± .004
0.976 ± .011
0.206
0.976
GA†
0.000 ± .000
0.000 ± .000
0.096
0.102
GD†
0.044 ± .007
0.598 ± .024
0.069
0.626
NPO
0.626 ± .071
0.968 ± .007
0.394
0.956
DUET
WGA
0.050 ± .002
0.996 ± .004
0.237
0.997
w/o MU
0.471
0.551
0.328
0.349
AdaPop
0.034 ± .014
0.948 ± .009
0.068
0.961
GA†
0.000 ± .000
0.000 ± .000
0.000
0.000
GD†
0.013 ± .003
0.334 ± .013
0.115
0.532
NPO
0.341 ± .013
0.773 ± .010
0.240
0.840
RWKU
WGA
0.040 ± .003
0.950 ± .007
0.135
0.965
Figure 4: Internal representation metrics at lr=10−4, across three models. Each row corresponds to one metric; each column corresponds to one data split. For the forget split, deeper erasure corresponds to lower ΔLP, higher ΔRank, lower Hid.Cos, and higher KL; directional desiderata are reversed for the retain split. Marker shape indicates benchmark (DUET vs. RWKU); colour indicates model family.
Table 4: Robustness at lr=10−4. Top: paraphrase ROUGE-L on the DUET forget split (lower is better). Bottom: RWKU Level-3 adversarial-attack ROUGE-L; lower indicates stronger resistance to knowledge recovery. Underline = best among NPO, WGA, AdaPop. GA and GD collapse and are excluded from forget-quality ranking.
Algo
Llama
Qwen
Gemma
DUET paraphrase forget ROUGE-L ↓
AdaPop
0.045
0.074
0.027
GA†
0.000
0.000
0.000
GD†
0.015
0.045
0.117
NPO
0.696
0.605
0.568
WGA
0.042
0.104
0.050
RWKU adversarial-attack forget ROUGE-L ↓
AdaPop
0.262
0.144
0.195
GA†
0.000
0.001
0.000
GD†
0.254
0.171
0.103
NPO
0.657
0.400
0.485
WGA
0.396
0.211
0.239
Figure 5: Component ablation on DUET. Top row: rare-fact paraphrase forget ROUGE-L (↓) and retain ROUGE-L (↑) over lr∈[2×10−5,8×10−4]. Bottom row: popular-fact paraphrase forget ROUGE-L (↓) and retain ROUGE-L (↑) over lr∈[2×10−4,8×10−3].
Table 5: DUET paraphrase forget ROUGE-L (↓) split by popularity tier at lr=10−4. Underline = best among NPO, WGA, AdaPop. GA is omitted (all values 0.000 under collapse); †GD’s low rare-tier values coincide with retain collapse (Table 1).
Algo
Llama
Qwen
Gemma
Pop.
Rare
Pop.
Rare
Pop.
Rare
AdaPop
0.040
0.049
0.055
0.092
0.028
0.026
GD†
0.027
0.003
0.086
0.003
0.232
0.002
NPO
0.854
0.538
0.778
0.432
0.796
0.339
WGA
0.067
0.017
0.194
0.014
0.096
0.003
Figure 6: ROUGE-L (solid) and Cosine Similarity (dashed) on the merged DUET forget (left) and retain (right) splits across learning rates, for Llama (top), Qwen (middle), and Gemma (bottom).
Table 6: Internal metrics at lr=10−4, averaged across Llama, Qwen, and Gemma. Directional arrows are shown in each section header. Underline = best among NPO, WGA, AdaPop. GA and GD values are reported for reference but indicate model collapse and are not ranked. Per-model breakdown in Appendix K.
Algo
ΔLP
ΔRank
Hid.Cos
KL
DUET — forget split (ΔLP↓, ΔRank↑, Hid.Cos↓, KL↑)
AdaPop
−160¯
+53,760¯
0.702
42.5
GA†
−3,106
+106,826
0.352
479.6
GD†
−5,917
+104,134
0.192
1006.4
NPO
−42
+8,824
0.861
5.4
WGA
−52
−3,714
0.729
16.2
RWKU — forget split (ΔLP↓, ΔRank↑, Hid.Cos↓, KL↑)
AdaPop
−23¯
+4,169¯
0.484
9.1
GA†
−6,083
+96,254
0.285
953.1
GD†
−6,842
+96,881
0.175
1139.7
NPO
−18
+2,739
0.883
2.9
WGA
−3.5
−11,010
0.541
6.7
Retain split — DUET (ΔLP↑, ΔRank↓, Hid.Cos↑, KL↓)
AdaPop
+21
−13,770
0.875
6.5
GA†
−2,750
+103,104
0.348
511.4
GD†
−251
−8,751
0.760
48.2
NPO
+21
−13,660
0.910
7.5
WGA
+23¯
−13,930¯
0.880
7.5
Retain split — RWKU (ΔLP↑, ΔRank↓, Hid.Cos↑, KL↓)
AdaPop
+33
−12,190
0.883
6.2
GA†
−6,106
+98,802
0.289
955.8
GD†
−540
−3,122
0.783
97.1
NPO
+31
−11,727
0.869
7.4
WGA
+34¯
−12,202¯
0.875
6.8
Figure 7: ROUGE-L (solid) and Cosine Similarity (dashed) on the RWKU forget (left) and retain (right) splits across learning rates, for Llama (top), Qwen (middle), and Gemma (bottom).
Table 7: MMLU accuracy and HellaSwag (HS) normalised accuracy at lr=10−4. w/o MU = the original checkpoint, with no unlearning applied. Values averaged across DUET and RWKU checkpoints.
Algo
Llama
Qwen
Gemma
MMLU
HS
MMLU
HS
MMLU
HS
w/o MU
0.65
0.73
0.71
0.68
0.47
0.64
AdaPop
0.65
0.76
0.71
0.68
0.52
0.62
GA
0.23
0.33
0.23
0.43
0.25
0.26
GD
0.64
0.59
0.71
0.65
0.47
0.37
NPO
0.65
0.72
0.70
0.68
0.51
0.63
WGA
0.66
0.76
0.71
0.70
0.52
0.68
Figure 8: DUET forget and retain splits separated by popularity tier (Llama). Top row: rare facts. Bottom row: popular facts. The order-of-magnitude difference in the learning rate required to achieve comparable forgetting across the two tiers is direct evidence of the popularity gap.
Table 8: AdaPop coefficient sensitivity. Each row perturbs a and/or b by ±20% from the analytically derived baseline. F. ↓ = forget ROUGE-L; R. ↑ = retain ROUGE-L. Δ values are signed differences from the baseline.
Config
a
b
F. ↓
R. ↑
ΔF.
ΔR.
baseline
58.70
0.7960
0.045
0.961
0
0
+a,+b
70.44
0.9552
0.038
0.975
−0.007
+0.014
+a,−b
70.44
0.6368
0.044
0.994
−0.001
+0.033
−a,+b
46.96
0.9552
0.024
0.920
−0.021
−0.041
−a,−b
46.96
0.6368
0.034
0.991
−0.011
+0.030
Figure 9: Forget (left) and retain (right) ROUGE-L Recall per training epoch on DUET, Llama-3.1-8B-Instruct, at lr=10−4.
Table 9: Discordant cases between Wikidata score and LLM-judge scores on DUET. ROUGE-L: Llama-3.1-8B-Instruct recall before unlearning. Top group: Wikidata-popular facts recalled at ROUGE-L 1.0 despite low LLM-judged salience, indicating deep memorisation. Bottom group: Wikidata-rare facts with higher LLM scores but variable recall, consistent with shallower encoding.
Question
Answer
Wikidata
LLM
ROUGE-L
Wikidata popular, LLM rare
What is the country of Ensenada?
Mexico
2113
30
1.0
What is the country of Tourcoing?
France
2112
40
1.0
What is the country of Qus?
Egypt
3763
50
1.0
Wikidata rare, LLM popular
What is the instance of Hua Hin?
seaside resort
148
9000
0.5
What is the instance of Weitra?
municipality of Austria
141
5000
0.7
What is the located in the administrative territorial entity of Ollantaytambo?
Urubamba Province
165
4000
1.0
Table 10: Agreement of each popularity proxy with measured corpus frequency on DUET, counted over Pile-train (383B tokens) with infini-gram. Pearson is computed in log space. The two proxies agree with each other at Spearman 0.596 and assign the same rare/popular label to 84% of facts.
Proxy
Pearson
Spearman
Wikidata score
0.970
0.839
LLM judge
0.677
0.659
Table 11: AdaPop under corrupted popularity scores (DUET, Llama-3.1-8B-Instruct, lr=10−4). Forget ROUGE-L is reported per popularity tier. Retain degrades by 0.04 even when every label is inverted.
Score noise
F. rare ↓
F. pop. ↓
R. ↑
none
0.05
0.04
0.96
log-normal, σ=1.0
0.02
0.05
0.96
labels 50% swapped
0.02
0.05
0.97
labels 100% swapped
0.01
0.05
0.92
Table 12: ROUGE-L Recall for additional baselines, Llama-3.1-8B-Instruct at lr=10−4. w/o MU = the original checkpoint, with no unlearning applied; AdaPop reproduced from Table 1 for reference. F. ↓ = forget; R. ↑ = retain. Underline = best among non-collapsed methods. †Model collapse.
Bench.
Algorithm
ROUGE-L
F. ↓
R. ↑
w/o MU
0.939
0.968
AdaPop
0.043
0.959
UNDIAL
0.884
0.998
RMU
0.882
0.998
PDU
0.476
0.954
NPO-SAM
0.684
0.970
SimNPO
0.340
0.999
SatImp
0.108
0.995
Adaptive RMU
0.933
0.963
AltPO
0.181
1.000
FLAT†
0.001
0.001
TPO
0.397
0.997
DUET
CE-U†
0.000
0.000
w/o MU
0.755
0.827
AdaPop
0.078
0.972
UNDIAL
0.568
0.939
RMU
0.801
0.986
PDU
0.201
0.909
NPO-SAM
0.529
0.885
SimNPO
0.573
0.987
SatImp
0.417
0.985
Adaptive RMU
0.162
0.819
AltPO
0.200
0.988
FLAT†
0.004
0.003
TPO
0.341
0.986
RWKU
CE-U†
0.000
0.000
Table 13: ROUGE-L Recall for additional baselines, Qwen2.5-7B-Instruct at lr=10−4. Notation as in Table 12.
Bench.
Algorithm
ROUGE-L
F. ↓
R. ↑
w/o MU
0.921
0.886
AdaPop
0.056
0.950
UNDIAL
0.797
0.930
RMU
0.734
0.983
PDU
0.632
0.869
NPO-SAM
0.211
0.769
SimNPO
0.295
0.991
SatImp
0.070
0.985
Adaptive RMU
0.562
0.831
AltPO
0.247
0.999
FLAT†
0.000
0.001
TPO
0.460
0.999
DUET
CE-U†
0.001
0.001
w/o MU
0.570
0.666
AdaPop
0.016
0.855
UNDIAL
0.313
0.668
RMU
0.407
0.922
PDU
0.455
0.710
NPO-SAM
0.298
0.596
SimNPO
0.338
0.934
SatImp
0.187
0.891
Adaptive RMU
0.118
0.466
AltPO
0.170
0.989
FLAT†
0.004
0.003
TPO
0.382
0.934
RWKU
CE-U†
0.003
0.001
Table 14: ROUGE-L Recall for additional baselines, Gemma-7B-it at lr=10−4. Notation as in Table 12.
Popular facts are memorised more deeply during pretraining and resist removal longer than rare ones, yet existing LLM unlearning methods apply uniform gradient pressure regardless of training-data frequency. We propose the AdaPop (Adaptive Popularity) method, which combines local token confidence with a per-fact popularity-dependent exponent derived from an external proxy (e.g., Wikidata sitelinks, LLM-as-Judge), and automates the forget-retain balance via a dual-ascent controller that adjusts the retain penalty each epoch. Across three model families and two benchmarks, AdaPop leaks ~5x less forgotten content than competing methods under paraphrased queries and ~1.6x less under adversarial reformulations. We support our analysis with internal metrics: under our method, forget-set hidden states move further from the pre-unlearning model's states than under other methods, while retain-set representations remain close.