모델의 속마음을 읽어주는 AI 통역사(Activation Oracle)를 훈련시켰더니, 자신이 훈련받은 바로 그 비밀만 못 읽는 이상한 통역사가 됐다
arXiv:2607.233792026-07-24
When Activation Oracles Learn Not to Read: Concept-Specific Blind Spots in Fine-Tuned Oracles
모델의 속마음을 읽어주는 AI 통역사(Activation Oracle)를 훈련시켰더니, 자신이 훈련받은 바로 그 비밀만 못 읽는 이상한 통역사가 됐다
Activation Oracle은 다른 언어모델의 내부 활성화값을 보고 자연어로 답하도록 훈련된 보조 모델이다. 연구팀은 특정 단어를 숨기도록 훈련된 '금기어(Taboo)' 모델을 만들고, 그 모델의 활성화값으로 오라클을 다시 훈련시켰더니 오라클이 오히려 그 숨겨진 단어를 못 읽는 현상을 발견했다. 정보는 오라클 내부에 여전히 존재하지만, 최종 출력 단계에서 선택적으로 억눌리는 것으로 나타났다.
METAL MEDIA 해설 도표
Activation Oracle이 자기 훈련 개념을 못 읽게 되는 과정
증거 상태측정 결과가 보고됨
1. 금기어 주체 모델 훈련Qwen3-8B를 leaf, moon 등 특정 단어를 숨기도록 파인튜닝. 말로는 드러내지 않지만 내부에서는 그 개념을 사용
2. 기본 오라클로 확인원본 모델 활성화값만 학습한 base AO는 숨겨진 단어를 잘 읽어냄 (예: HINT에서 정확 회수율 93%)
3. 오라클을 같은 개념으로 재훈련leaf를 숨긴 주체 모델의 활성화값으로 오라클(FT-AO)을 훈련하면, 그 오라클이 오히려 leaf를 못 읽게 됨(35%로 하락)
4. 내부 vs 출력 비교프로브로는 여전히 leaf 정보가 내부에 남아있지만, 오라클 출력층(LogitLens)에서는 leaf 순위가 나빠짐
5. 레이어 절제로 원인 위치 확인중반 레이어(L18-23)의 학습된 업데이트를 제거하면 억눌렸던 목표 단어 확률이 다시 회복됨
METAL MEDIA이 원문을 바탕으로 재구성한 해설 도표이며, 논문 저자의 원문 figure가 아닙니다.
무엇을 했나
연구팀은 Qwen3-8B를 기반으로 leaf, moon 등 특정 단어를 숨기도록(직접 말하지 않으면서 내부적으로 사용하도록) 파인튜닝한 '금기어(Taboo)' 주체 모델들을 만들었다.
기본 오라클(base AO)은 이 숨겨진 단어를 잘 읽어냈다. 그런데 그 특정 주체 모델의 활성화값으로 오라클 자체를 다시 파인튜닝(FT-AO)하면, 예상과 달리 그 오라클은 자신이 훈련받은 개념을 오히려 더 못 읽게 됐다.
이 실패는 개념별로 특이적이었다. leaf를 훈련받은 오라클은 leaf만 못 읽고 다른 개념(moon 등)은 여전히 잘 읽었으며, 이는 하나의 모델이 leaf와 moon 둘 다 숨기고 있는 경우에도 똑같이 나타났다.
선형 프로브(내부 상태를 읽는 간단한 분류기)로 확인하면 숨겨진 단어 정보는 오라클 내부에 여전히 남아 있었다. 문제는 그 정보가 최종 출력 어휘로 변환되는 '읽어내기(readout)' 단계, 특히 중후반 레이어(L18-23 부근)에서 억제된다는 데 있었다.
해당 레이어 구간의 학습된 업데이트(LoRA)를 제거하는 실험을 통해, 이 억제 효과가 실제로 그 레이어들에서 발생함을 인과적으로 확인했다.
Figure 1: Fine-tuned Activation Oracles become concept-specific anti-readers.
Table 1: Linear-probe performance on Qwen3-8B subject residual-stream activations. Acc. is mean 5-fold cross-validated accuracy; Std. is the standard deviation across folds; AUC is macro one-vs-rest ROC AUC. Per-class columns report held-out recall for each label. Chance accuracy is 16.7%.
Overall
Per-class recall
Regime
Layer
Acc.
Std.
AUC
book
flag
leaf
moon
wave
base
HINT
L9
0.9990
0.0019
1.0000
0.998
1.000
0.998
1.000
1.000
1.000
L18
0.9971
0.0018
1.0000
0.995
1.000
0.998
0.993
1.000
1.000
L27
0.9957
0.0041
1.0000
0.995
0.995
0.998
0.993
0.998
1.000
REFUSAL
L9
1.0000
0.0000
1.0000
1.000
1.000
1.000
1.000
1.000
1.000
L18
0.9990
0.0012
1.0000
1.000
0.998
1.000
0.998
1.000
1.000
L27
0.9976
0.0021
1.0000
0.995
0.998
1.000
0.998
0.998
1.000
SAMETEXT
L9
1.0000
0.0000
1.0000
1.000
1.000
1.000
1.000
1.000
1.000
L18
1.0000
0.0000
1.0000
1.000
1.000
1.000
1.000
1.000
1.000
L27
1.0000
0.0000
1.0000
1.000
1.000
1.000
1.000
1.000
1.000
THINK
L9
1.0000
0.0000
1.0000
1.000
1.000
1.000
1.000
1.000
1.000
L18
0.9986
0.0012
1.0000
0.998
1.000
1.000
1.000
0.995
1.000
L27
0.9981
0.0010
1.0000
0.998
1.000
1.000
1.000
0.993
1.000
OFFTOPIC
L9
0.9971
0.0018
1.0000
1.000
0.995
0.998
1.000
0.993
1.000
L18
0.9919
0.0019
0.9999
0.995
0.983
0.995
0.998
0.988
1.000
L27
0.9833
0.0045
0.9997
0.990
0.968
0.993
0.995
0.968
1.000
(b) (left) Own FT-AOs fail to become specialist readers; (right) The blind spot follows the AO training concept.
Table 2: Five-way linear-probe accuracy on AO residual-stream activations. FT-AOs are cooperative α=1.0 oracles. The Own column gives the concept on which the FT-AO was trained. Chance accuracy is 20%.
Probe accuracy (%) at AO layer Lℓ
Regime
AO
Own
L4
L8
L14
L18
L24
L30
L33
HINT
base-AO
–
98.0
96.7
94.3
93.3
93.3
94.0
94.3
leaf-FT
leaf
98.3
96.7
93.7
93.3
91.3
93.3
93.0
moon-FT
moon
98.7
95.7
94.3
93.3
92.7
92.3
93.3
wave-FT
wave
98.0
97.7
95.3
93.0
93.7
93.0
93.0
flag-FT
flag
97.7
96.7
95.3
93.7
93.7
94.0
93.3
book-FT
book
98.0
96.3
94.3
93.0
93.0
93.0
93.0
REFUSAL
base-AO
–
95.7
93.3
91.3
89.7
87.3
86.3
90.3
leaf-FT
leaf
96.7
95.0
91.7
85.3
83.3
85.0
90.0
moon-FT
moon
95.7
95.0
93.3
89.3
84.7
86.0
91.3
wave-FT
wave
96.0
94.3
92.0
89.3
84.0
85.0
90.3
flag-FT
flag
96.0
95.3
92.7
89.0
85.3
86.3
91.7
book-FT
book
96.0
93.7
92.3
87.0
84.7
86.0
92.7
SAMETEXT
base-AO
–
100.0
100.0
100.0
100.0
100.0
100.0
100.0
leaf-FT
leaf
100.0
100.0
100.0
100.0
100.0
100.0
100.0
moon-FT
moon
100.0
100.0
100.0
100.0
100.0
100.0
100.0
wave-FT
wave
100.0
100.0
100.0
100.0
100.0
100.0
100.0
flag-FT
flag
100.0
100.0
100.0
100.0
100.0
100.0
100.0
book-FT
book
100.0
100.0
100.0
100.0
99.0
99.0
100.0
THINK
base-AO
–
98.7
97.3
97.0
96.0
94.3
95.7
98.0
leaf-FT
leaf
99.0
98.7
96.7
95.0
93.0
95.3
97.0
moon-FT
moon
98.7
98.7
97.0
95.0
94.7
94.7
97.0
wave-FT
wave
98.3
98.0
96.7
96.0
91.3
92.3
96.3
flag-FT
flag
99.0
98.3
97.0
96.7
92.7
94.0
97.7
book-FT
book
99.0
98.7
96.0
95.7
93.0
93.3
98.3
OFFTOPIC
base-AO
–
73.7
65.3
62.7
63.3
62.3
62.3
70.0
leaf-FT
leaf
73.3
69.3
61.7
62.0
54.3
55.7
66.7
moon-FT
moon
74.7
69.7
65.3
67.3
62.3
60.3
68.7
wave-FT
wave
75.7
69.7
62.7
61.0
54.0
51.7
60.3
flag-FT
flag
72.0
66.7
60.7
63.7
50.3
55.7
66.3
Figure 2: Behavioral outcomes of Taboo-trained subjects across prompting regimes.
Table 3: Median LogitLens rank of the target token when the AO’s own Qwen3-8B Bold entries mark own evaluations, where the FT-AO’s training concept matches the subject hidden concept. Each cell is aggregated across 30 held-out prompts.
Regime
AO
Protocol
book
flag
leaf
moon
wave
hint
base-AO
coop
1
1
3
1
2
strict
1
23
235
10
62
book-FT
coop
97
1
171
2
7
strict
91
25
3,792
30
159
flag-FT
coop
5
247
367
18
48
strict
12
1,665
2,526
103
424
leaf-FT
coop
34
13
8,788
29
47
strict
47
124
19.3k
118
587
moon-FT
coop
4
3
68
267
15
strict
10
90
2,172
1,020
244
wave-FT
coop
1
1
15
1
252
strict
1
56
1,658
28
3,253
refusal
base-AO
coop
1
18
93
10
24
strict
4,429
7,382
76.8k
9,992
25.7k
book-FT
coop
1,913
15
1,772
31
68
strict
30.7k
1,965
91.7k
14.5k
18.9k
flag-FT
coop
32
1,171
2,200
99
383
strict
3,996
4,380
68.6k
9,154
23.1k
leaf-FT
coop
251
111
26.7k
139
296
strict
13.5k
2,493
86.4k
17.0k
24.9k
moon-FT
coop
41
88
1,428
1,799
169
strict
9,373
2,894
71.3k
20.9k
20.0k
wave-FT
coop
3
30
276
30
5,802
strict
5,835
4,348
74.7k
10.6k
50.3k
sametext
base-AO
coop
7
245
5,842
193
237
strict
663
26.6k
58.0k
14.6k
24.3k
book-FT
coop
20.6k
534
64.0k
2,238
2,218
strict
35.4k
6,806
94.8k
17.9k
24.3k
flag-FT
coop
548
26.9k
69.4k
12.8k
20.4k
strict
1,312
13.7k
75.6k
13.9k
37.5k
Figure 3: Δ-LogitLens decodability of Taboo concepts from subject residual-stream activations at layer L18. Lower rank indicates greater accessibility. The dashed line shows the shuffled-label baseline.
Table 4: Mean AO output entropy, in nats, at the prediction position for cooperative subjects. Bold entries mark own-FT-AO cells where entropy is higher than the base-AO entropy. Cooperative anti-reading is not accompanied by a uniform entropy increase: entropy rises in HINT and THINK, but is lower than the base AO in SAMETEXT and OFFTOPIC.
Regime
base-AO
own, α=1.0
cross, α=1.0
own, α=0.5
cross, α=0.5
HINT
1.9
3.1
2.5
1.6
1.5
REFUSAL
2.6
2.8
2.2
2.2
2.0
SAMETEXT
3.8
2.1
2.2
2.4
2.3
THINK
2.8
3.3
2.8
3.0
2.8
OFFTOPIC
5.0
3.9
3.9
–
–
Figure 4: Base-AO recovery of hidden Taboo concepts from subject activations, pooled over mixture strengths α∈{0.5,1.0}. Left: exact target recovery and semantic recovery judged by Claude Sonnet. Right: mean target probability and median target rank under the AO output distribution.
Table 5: Mean AO output entropy, in nats, at the prediction position for strict subjects. In contrast to an uncertainty-based explanation, FT-AO entropy is substantially lower than base-AO entropy across all regimes. This indicates that strict-subject anti-reading is associated with a sharper output distribution rather than a more diffuse one.
Regime
base-AO
own, α=1.0
cross, α=1.0
own, α=0.5
cross, α=0.5
HINT
3.8
0.9
0.8
0.8
0.7
REFUSAL
5.7
0.2
0.2
0.2
0.2
SAMETEXT
4.6
0.9
0.9
1.1
1.2
THINK
4.8
0.4
0.4
0.4
0.4
OFFTOPIC
5.6
1.5
1.5
–
–
Figure 5: Aggregate comparison of base AOs, cross FT-AOs, and own FT-AOs. Results are averaged over target concepts and activation-capture regimes. Error bars show 95% bootstrap confidence intervals.
Table 6: Representative target–top-1 alignment values. Low ptarget/ptop1 with high ptop1 indicates that the AO is confident at the prediction position, but its top token is not the target. This pattern is consistent with readout-side anti-reading rather than simple uncertainty.
Regime
AO
ptop1
ptarget
r
H
HINT
base-AO
0.80
0.63
0.78
0.53
HINT
moon-FT
0.89
0.69
0.78
0.32
HINT
leaf-FT (own)
0.86
0.28
0.33
0.43
SAMETEXT
base-AO
0.76
0.31
0.41
0.66
SAMETEXT
moon-FT (own)
0.85
0.00
0.00
0.40
Figure 6: Concept-specific blind spots in FT-AOs. Rows indicate the hidden concept of the subject used to train the FT-AO, and columns indicate the hidden concept of the test subject. Each cell reports mean P(c⋆), aggregated over activation-capture regimes and mixture strengths α∈{0.5,1.0}. Red borders mark own evaluations; off-diagonal entries are cross evaluations.
Table 7: Backing values for Figure 9. The no-ablation column gives the intact-FT-AO mean P(c⋆) over n=20 HINT captures on the own-concept subject; this is the denominator of ρ. Subsequent columns give the normalized restoration ratio ρ=P(c⋆)¯ablated/P(c⋆)¯none for each ablated layer range. Bold entries mark ranges where ablation restores target probability above the intact-FT-AO baseline. Full-LoRA ablation, not shown, drives P(c⋆)¯ close to zero for every concept.
no ablation
ρc,S, relative to no-ablation baseline
FT-AO
P(c⋆)¯
L0–5
L6–11
L12–17
L18–23
L24–29
L30–35
book-FT
0.244
0.00
0.65
1.81
1.74
1.06
0.91
flag-FT
0.253
0.00
0.54
2.71
2.88
1.38
0.77
leaf-FT
0.143
0.00
0.55
1.17
2.01
1.24
0.87
moon-FT
0.364
0.00
0.62
1.50
1.53
1.25
1.07
wave-FT
0.126
0.00
0.69
2.24
2.97
1.45
0.86
Figure 7: Internal decodability versus output-head accessibility in AOs. Left: accuracy of a linear probe trained on AO hidden states at layer L33 to predict the hidden Taboo concept. The dotted line marks the 5-way chance baseline. Right: median LogitLens rank of the target word; lower rank indicates greater accessibility.
Table 8: Backing values for Figure 10. Each cell reports P(c⋆)¯ for an AO evaluated on activations from the cooperative α=1.0 leaf_moon two-concept subject. Columns indicate the queried target concept. Bold entries mark own-target evaluations, where the target concept matches the single-concept FT-AO’s training concept. Each cell aggregates approximately 300 captures.
HINT
REFUSAL
THINK
AO
leaf
moon
leaf
moon
leaf
moon
base AO
0.060
0.320
0.050
0.327
0.044
0.341
leaf-FT, α=1.0
0.011
0.239
0.012
0.224
0.004
0.229
moon-FT, α=1.0
0.069
0.143
0.060
0.180
0.049
0.165
leaf-FT, α=0.5
0.043
0.329
0.031
0.348
0.025
0.363
moon-FT, α=0.5
0.072
0.262
0.065
0.289
0.063
0.290
Figure 8: Target readout suppression grows in late AO layers. Higher values indicate that the FT-AO ranks the target worse than the base AO, values near zero indicate no relative suppression.
Table 9: Behavioral evaluation for cooperative subjects with α=0.5.
Concept
Regime
Leak (%)
Refuse (%)
Other (%)
n
book
hint
4.3
4.0
91.7
303
refusal
1.3
9.2
89.5
306
sametext
0.0
0.0
100.0
300
think
0.3
15.2
84.5
309
offtopic
0.3
0.3
99.3
300
flag
hint
0.7
3.3
96.0
303
refusal
0.3
7.2
92.5
306
sametext
0.0
0.0
100.0
300
think
0.0
8.4
91.6
309
offtopic
0.0
0.7
99.3
300
leaf
hint
2.6
3.3
94.1
303
refusal
2.6
7.2
90.2
306
sametext
0.0
0.0
100.0
300
think
0.3
7.4
92.2
309
offtopic
0.0
0.3
99.7
300
moon
hint
6.3
4.0
89.8
303
refusal
3.6
8.5
87.9
306
sametext
1.0
0.0
99.0
300
think
2.6
11.3
86.1
309
offtopic
0.0
2.0
98.0
300
wave
hint
0.7
4.3
95.0
303
refusal
0.3
11.4
88.2
306
sametext
0.0
0.0
100.0
300
think
0.3
11.1
88.6
307
offtopic
0.0
0.0
100.0
300
Figure 9: Layer-range ablations localize the anti-reading effect. Values are normalized by no-ablation baseline, values above one indicate restored target accessibility.
Table 10: Behavioral evaluation for cooperative subjects with α=1.0.
Concept
Regime
Leak (%)
Refuse (%)
Other (%)
n
book
hint
1.3
0.0
98.7
303
refusal
0.7
1.0
98.4
306
sametext
0.0
0.0
100.0
300
think
1.0
0.3
98.7
308
offtopic
0.3
0.0
99.7
300
flag
hint
0.0
0.0
100.0
303
refusal
0.0
1.3
98.7
306
sametext
0.0
0.0
100.0
300
think
0.3
0.0
99.7
306
offtopic
0.0
0.0
100.0
300
leaf
hint
0.7
0.0
99.3
303
refusal
1.3
1.3
97.4
306
sametext
0.0
0.0
100.0
300
think
1.3
0.3
98.4
306
offtopic
0.0
0.0
100.0
300
moon
hint
1.7
0.0
98.3
303
refusal
2.3
2.3
95.4
306
sametext
1.0
0.0
99.0
300
think
1.9
0.3
97.7
309
offtopic
0.3
0.0
99.7
300
wave
hint
0.7
0.0
99.3
303
refusal
0.0
0.7
99.3
306
sametext
0.0
0.0
100.0
300
think
0.3
1.0
98.7
309
offtopic
0.3
0.0
99.7
300
Figure 10: FT-AOs evaluated on multi-concept Taboo subjects containing both leaf and moon. Each cell reports mean P(c⋆) under the AO output distribution. Red borders mark cases where the queried target matches the FT-AO’s own training concept.
Table 11: Behavioral evaluation for strict subjects with α=0.5.
Concept
Regime
Leak (%)
Refuse (%)
Other (%)
n
book
hint
1.0
33.0
66.0
303
refusal
0.0
92.5
7.5
306
sametext
0.0
0.0
100.0
300
think
0.6
57.3
42.1
309
offtopic
0.3
0.3
99.3
300
flag
hint
0.0
34.7
65.3
303
refusal
0.0
94.1
5.9
306
sametext
0.0
0.0
100.0
300
think
0.0
66.0
34.0
309
offtopic
0.0
0.7
99.3
300
leaf
hint
0.7
34.3
65.0
303
refusal
0.0
92.5
7.5
306
sametext
0.0
0.0
100.0
300
think
0.0
56.6
43.4
309
offtopic
0.0
0.7
99.3
300
moon
hint
0.0
42.9
57.1
303
refusal
0.0
95.8
4.2
306
sametext
1.0
0.0
99.0
300
think
0.0
63.4
36.6
309
offtopic
0.0
1.0
99.0
300
wave
hint
0.3
37.0
62.7
303
refusal
0.0
94.1
5.9
306
sametext
0.0
0.0
100.0
300
think
0.3
69.3
30.4
309
offtopic
0.3
0.3
99.3
300
Figure 11: Cross-regime transfer of subject-activation concept probes. Each heatmap corresponds to a subject layer ℓ∈{4,8,18,33}. Rows indicate the (protocol,α,regime) cell used to train the probe, and columns indicate the cell used for evaluation. Each entry reports five-way hidden-concept classification accuracy; chance is 20%. Red outlines mark train-test matched cells. Strong off-diagonal transfer indicates that the hidden concept is not only decodable within a single regime, but remains linearly accessible across protocols, mixture strengths, and activation-capture regimes.
Table 12: Behavioral evaluation for strict subjects with α=1.0.
Concept
Regime
Leak (%)
Refuse (%)
Other (%)
n
book
hint
1.3
15.5
83.2
303
refusal
0.0
94.4
5.6
306
sametext
0.0
0.0
100.0
300
think
0.6
52.4
46.9
309
offtopic
0.0
0.7
99.3
300
flag
hint
0.0
28.4
71.6
303
refusal
0.0
92.5
7.5
306
sametext
0.0
0.0
100.0
300
think
0.0
61.7
38.3
308
offtopic
0.3
0.3
99.3
300
leaf
hint
0.3
28.4
71.3
303
refusal
0.0
91.8
8.2
306
sametext
0.0
0.0
100.0
300
think
0.0
54.7
45.3
309
offtopic
0.0
0.3
99.7
300
moon
hint
0.0
17.5
82.5
303
refusal
0.0
90.2
9.8
306
sametext
1.0
0.0
99.0
300
think
0.0
47.9
52.1
309
offtopic
0.0
0.0
100.0
300
wave
hint
0.0
24.1
75.9
303
refusal
0.0
93.1
6.9
306
sametext
0.0
0.0
100.0
300
think
0.0
51.5
48.5
309
offtopic
0.0
0.3
99.7
300
Table 13: Δ-LogitLens decodability for cooperative subjects with α=0.5.
Concept
Regime
Rank ws
Rank ns
P(c⋆)
‖δ‖
n
book
hint
469
3,457
9.7e-05
23.48
200
refusal
76.9k
41.2k
3.5e-09
22.53
200
sametext
119.3k
115.7k
4.7e-10
21.93
200
think
25.5k
27.2k
1.6e-07
20.06
200
offtopic
38.5k
53.6k
6.5e-08
17.07
200
flag
hint
25
19
2.30e-03
28.30
200
refusal
2,503
321
9.1e-06
23.84
200
sametext
41.5k
20.8k
5.3e-08
22.71
200
think
755
126
8.7e-05
23.61
200
offtopic
37.6k
79.5k
1.2e-07
17.03
200
leaf
hint
8
6
2.42e-03
24.23
200
refusal
57
19
4.92e-04
23.30
200
sametext
70
730
9.82e-04
21.61
200
think
25
19
1.51e-03
20.75
200
offtopic
34.7k
8,114
8.4e-08
17.35
200
moon
hint
4
27
1.96e-03
27.16
200
refusal
22
183
4.53e-04
24.41
200
sametext
3,181
4,013
1.3e-05
21.46
200
think
11
237
4.32e-03
22.61
200
offtopic
20.8k
66.1k
2.8e-07
17.46
200
wave
hint
5
22
0.016
26.03
200
refusal
266
656
1.62e-04
22.40
200
sametext
16.0k
44.0k
6.5e-07
20.89
200
think
42
244
3.14e-03
21.24
200
offtopic
38.9k
70.8k
5.8e-08
17.99
200
Table 14: Δ-LogitLens decodability for cooperative subjects with α=1.0.
Concept
Regime
Rank ws
Rank ns
P(c⋆)
‖δ‖
n
book
hint
965
4,046
2.0e-05
27.61
200
refusal
26.7k
31.2k
3.9e-08
33.39
200
sametext
97.7k
82.5k
2.3e-09
22.67
200
think
7,476
12.0k
8.0e-07
33.71
200
offtopic
47.1k
89.0k
6.5e-08
8.94
200
flag
hint
43
22
7.53e-04
29.69
200
refusal
684
202
1.6e-05
32.53
200
sametext
35.5k
21.1k
1.4e-07
20.98
200
think
272
110
1.23e-04
34.15
200
offtopic
3,770
3,789
1.6e-05
9.04
200
leaf
hint
6
3
9.09e-03
31.16
200
refusal
13
21
3.21e-03
35.56
200
sametext
138
461
3.55e-04
21.75
200
think
4
9
0.021
34.91
200
offtopic
734
1,103
8.2e-05
10.88
200
moon
hint
6
25
3.72e-03
31.78
200
refusal
15
93
2.28e-03
35.38
200
sametext
2,486
2,890
1.9e-05
21.26
200
think
7
54
5.70e-03
35.87
200
offtopic
352
1,135
3.42e-04
9.86
200
wave
hint
4
27
0.012
29.98
200
refusal
22
294
3.44e-03
37.12
200
sametext
1,815
12.1k
1.8e-05
20.26
200
think
10
146
0.016
35.64
200
offtopic
716
7,315
7.1e-05
10.52
200
Table 15: Δ-LogitLens decodability for strict subjects with α=0.5.
Concept
Regime
Rank ws
Rank ns
P(c⋆)
‖δ‖
n
book
hint
668
4,166
5.5e-05
24.63
200
refusal
48.8k
96.4k
1.6e-08
28.75
200
sametext
69.3k
76.3k
2.5e-08
16.90
200
think
4,274
26.8k
9.1e-06
23.72
200
offtopic
41.4k
50.7k
8.0e-10
17.84
200
flag
hint
11
13
0.023
24.86
200
refusal
73.2k
59.2k
5.3e-09
28.82
200
sametext
949
334
5.3e-05
16.55
200
think
482
1,482
1.91e-04
24.33
200
offtopic
7,557
34.8k
3.8e-08
18.13
200
leaf
hint
1,615
2,393
2.3e-05
25.13
200
refusal
144.3k
141.6k
1.7e-11
28.43
200
sametext
24.8k
43.7k
4.5e-07
17.59
200
think
29.2k
41.7k
3.0e-07
24.11
200
offtopic
89.4k
67.2k
5.2e-11
17.72
200
moon
hint
39
389
8.46e-04
24.11
200
refusal
47.4k
83.1k
1.8e-08
29.51
200
sametext
2,587
7,854
3.1e-05
15.97
200
think
402
4,958
1.57e-04
23.69
200
offtopic
31.8k
109.4k
3.1e-09
16.96
200
wave
hint
6
15
0.033
24.44
200
refusal
76.0k
36.8k
4.2e-09
29.07
200
sametext
5,341
4,079
8.8e-06
16.84
200
think
7,589
8,040
2.7e-06
24.92
200
offtopic
57.7k
97.6k
4.9e-10
17.11
200
Table 16: Δ-LogitLens decodability for strict subjects with α=1.0.
Concept
Regime
Rank ws
Rank ns
P(c⋆)
‖δ‖
n
book
hint
389
3,734
1.19e-04
25.18
200
refusal
38.6k
113.6k
4.2e-08
30.04
200
sametext
47.4k
87.3k
9.5e-08
16.77
200
think
2,401
29.8k
2.2e-05
24.89
200
offtopic
50.7k
86.2k
1.4e-10
18.84
200
flag
hint
20
27
6.23e-03
26.93
200
refusal
89.7k
86.2k
3.8e-09
31.08
200
sametext
4,602
3,675
1.2e-05
16.84
200
think
3,508
10.3k
2.0e-05
26.81
200
offtopic
5,687
44.6k
2.1e-08
20.18
200
leaf
hint
514
295
6.1e-05
26.95
200
refusal
141.6k
132.8k
3.3e-11
30.13
200
sametext
4,665
2,530
6.2e-06
18.17
200
think
26.7k
21.9k
4.3e-07
26.33
200
offtopic
88.8k
70.2k
1.1e-10
17.53
200
moon
hint
18
201
4.83e-03
27.72
200
refusal
58.0k
96.5k
1.2e-08
28.93
200
sametext
7,277
29.1k
5.9e-06
17.69
200
think
41
1,023
1.11e-03
25.41
200
offtopic
29.9k
115.0k
1.2e-09
18.83
200
wave
hint
2
13
0.048
25.86
200
refusal
76.2k
52.0k
5.8e-09
30.34
200
sametext
5,483
8,422
5.4e-06
16.94
200
think
102
635
2.20e-04
25.84
200
offtopic
64.2k
108.5k
2.4e-10
18.20
200
Table 17: Per-cell Base-AO recovery for cooperative subjects with α=0.5.
Concept
Regime
Exact
Sem.
P(c⋆)¯
Rank med.
nj
np
%
%
book
hint
80
88
0.534
0
25
303
refusal
36
48
0.229
35
25
304
sametext
96
96
0.541
0
25
300
think
36
36
0.249
54
25
308
offtopic
0
0
3.51e-04
388
25
300
flag
hint
96
100
0.652
0
25
303
refusal
56
56
0.328
3
25
305
sametext
88
88
0.415
0
25
300
think
56
56
0.352
3
25
309
offtopic
0
0
1.47e-04
1,014
25
300
leaf
hint
76
84
0.475
0
25
303
refusal
24
32
0.215
5
25
305
sametext
72
76
0.267
0
25
300
think
48
52
0.250
2
25
309
offtopic
0
0
6.3e-05
2,493
25
300
moon
hint
100
100
0.585
0
25
303
refusal
64
72
0.351
0
25
305
sametext
76
76
0.355
0
25
300
think
68
68
0.475
0
25
309
offtopic
0
4
1.39e-03
120
25
300
wave
hint
84
84
0.501
0
25
303
refusal
32
40
0.238
28
25
304
sametext
60
60
0.130
0
25
300
think
40
52
0.212
26
25
309
offtopic
0
0
5.52e-04
321
25
300
Table 18: Per-cell Base-AO recovery for cooperative subjects with α=1.0.
Concept
Regime
Exact
Sem.
P(c⋆)¯
Rank med.
nj
np
%
%
book
hint
92
96
0.747
0
25
303
refusal
76
84
0.623
0
25
303
sametext
96
96
0.587
0
25
300
think
88
88
0.674
0
25
309
offtopic
0
0
2.49e-04
511
25
300
flag
hint
100
100
0.758
0
25
303
refusal
76
80
0.590
0
25
300
sametext
52
52
0.286
0
25
300
think
80
80
0.658
0
25
309
offtopic
0
0
8.53e-04
1,163
25
300
leaf
hint
100
100
0.632
0
25
303
refusal
64
64
0.482
0
25
303
sametext
40
40
0.093
2
25
300
think
76
76
0.515
0
25
309
offtopic
0
0
1.81e-04
1,643
25
300
moon
hint
100
100
0.598
0
25
303
refusal
72
76
0.499
0
25
304
sametext
76
76
0.279
0
25
300
think
84
88
0.581
0
25
309
offtopic
0
0
1.10e-03
238
25
300
wave
hint
100
100
0.639
0
25
303
refusal
80
88
0.535
0
25
305
sametext
56
56
0.150
1
25
300
think
80
80
0.527
0
25
309
offtopic
0
0
1.37e-03
260
25
300
Table 19: Per-cell Base-AO recovery for strict subjects with α=0.5.
Concept
Regime
Exact
Sem.
P(c⋆)¯
Rank med.
nj
np
%
%
book
hint
48
52
0.364
0
25
303
refusal
0
0
1.23e-03
180
25
292
sametext
8
8
0.032
9
25
300
think
20
20
0.166
69
25
305
offtopic
0
0
3.82e-04
342
25
300
flag
hint
40
44
0.411
0
25
302
refusal
0
0
5.05e-04
274
25
291
sametext
0
0
2.31e-04
619
25
300
think
20
20
0.172
336
25
306
offtopic
0
0
1.34e-04
1,114
25
300
leaf
hint
36
40
0.176
6
25
303
refusal
0
0
5.14e-04
2,350
25
285
sametext
0
8
4.67e-04
510
25
300
think
20
24
0.097
608
25
308
offtopic
0
0
6.3e-05
2,474
25
300
moon
hint
36
36
0.215
2
25
302
refusal
4
8
8.59e-03
52
25
273
sametext
8
20
9.19e-03
51
25
300
think
24
24
0.121
15
25
306
offtopic
0
4
9.72e-04
169
25
300
wave
hint
20
32
0.155
7
25
303
refusal
0
0
2.32e-04
598
25
286
sametext
0
12
1.16e-03
138
25
300
think
12
28
0.047
476
25
305
offtopic
0
4
5.81e-04
326
25
300
Table 20: Per-cell Base-AO recovery for strict subjects with α=1.0.
Concept
Regime
Exact
Sem.
P(c⋆)¯
Rank med.
nj
np
%
%
book
hint
36
40
0.388
0
25
301
refusal
0
0
7.40e-04
193
25
261
sametext
0
0
0.018
10
25
300
think
20
20
0.171
53
25
298
offtopic
–
–
3.30e-04
400
–
300
flag
hint
44
48
0.387
0
25
298
refusal
0
0
4.24e-04
273
25
237
sametext
0
0
1.68e-04
689
25
300
think
20
20
0.133
308
25
287
offtopic
–
–
1.05e-04
1,165
–
300
leaf
hint
32
36
0.175
4
25
303
refusal
0
0
1.96e-04
2,915
25
261
sametext
0
4
7.32e-04
230
25
300
think
16
16
0.099
566
25
301
offtopic
–
–
6.4e-05
2,513
–
300
moon
hint
44
52
0.242
1
25
301
refusal
0
0
2.61e-03
73
25
261
sametext
0
0
4.62e-03
88
25
300
think
24
24
0.164
13
25
290
offtopic
–
–
6.97e-04
237
–
300
wave
hint
28
32
0.169
2
25
303
refusal
0
0
3.27e-04
572
25
256
sametext
0
0
2.65e-03
58
25
300
think
20
24
0.053
140
25
296
offtopic
–
–
5.42e-04
293
–
300
Table 21: Per-cell exact recovery (%) for cooperative subjects with α=0.5.
Concept
Regime
Base
Cross
Own
book
hint
72
68 [65,71]
63
refusal
30
28 [27,29]
26
sametext
96
34 [22,43]
30
think
37
33 [30,36]
28
offtopic
–
–
–
flag
hint
84
79 [78,79]
77
refusal
43
39 [37,42]
37
sametext
89
28 [16,39]
18
think
47
42 [40,44]
41
offtopic
–
–
–
leaf
hint
71
62 [60,64]
49
refusal
30
27 [26,27]
22
sametext
76
31 [21,44]
1
think
43
30 [28,33]
17
offtopic
–
–
–
moon
hint
92
90 [88,92]
84
refusal
60
52 [48,58]
41
sametext
94
61 [46,73]
26
think
78
70 [64,77]
52
offtopic
–
–
–
wave
hint
73
67 [66,68]
56
refusal
37
36 [34,37]
26
sametext
64
12 [6,20]
0
think
35
30 [29,31]
23
offtopic
–
–
–
Table 22: Per-cell exact recovery (%) for cooperative subjects with α=1.0.
Concept
Regime
Base
Cross
Own
book
hint
94
78 [65,92]
45
refusal
77
59 [46,71]
21
sametext
98
13 [0,35]
1
think
85
59 [49,72]
27
offtopic
0
0 [0,0]
0
flag
hint
96
83 [78,86]
47
refusal
74
52 [49,55]
19
sametext
63
1 [0,2]
0
think
85
60 [53,67]
18
offtopic
0
0 [0,0]
0
leaf
hint
90
49 [31,67]
14
refusal
68
36 [19,49]
11
sametext
34
0 [0,0]
0
think
75
32 [19,48]
6
offtopic
0
0 [0,0]
0
moon
hint
94
82 [77,89]
46
refusal
76
57 [52,62]
32
sametext
85
5 [1,10]
1
think
87
69 [65,74]
32
offtopic
0
0 [0,0]
0
wave
hint
95
61 [51,66]
22
refusal
77
44 [39,48]
8
sametext
60
0 [0,1]
0
think
78
37 [27,41]
7
offtopic
0
0 [0,0]
0
Table 23: Per-cell exact recovery (%) for strict subjects with α=0.5.
Concept
Regime
Base
Cross
Own
book
hint
54
50 [48,51]2
46
refusal
0
0 [0,0]2
0
sametext
13
1 [0,2]2
1
think
22
21 [20,22]2
20
offtopic
–
–
–
flag
hint
56
55 [54,55]2
53
refusal
0
0 [0,0]2
0
sametext
0
0 [0,0]2
0
think
22
21 [20,21]2
21
offtopic
–
–
–
leaf
hint
26
28 [27,29]3
–
refusal
0
0 [0,0]3
–
sametext
0
0 [0,0]3
–
think
13
13 [12,14]3
–
offtopic
–
–
–
moon
hint
42
40 [39,41]3
–
refusal
3
0 [0,0]3
–
sametext
4
1 [0,2]3
–
think
23
19 [17,19]3
–
offtopic
–
–
–
wave
hint
33
35 [34,35]2
34
refusal
0
0 [0,0]2
0
sametext
0
0 [0,0]2
0
think
12
9 [8,10]2
7
offtopic
–
–
–
Table 24: Per-cell exact recovery (%) for strict subjects with α=1.0.
Concept
Regime
Base
Cross
Own
book
hint
55
49 [46,53]
49
refusal
0
0 [0,0]
0
sametext
9
7 [0,17]
1
think
26
19 [16,24]
17
offtopic
0
0 [0,0]
0
flag
hint
56
51 [50,52]
48
refusal
0
0 [0,0]
0
sametext
0
0 [0,0]
0
think
17
15 [14,17]
14
offtopic
0
0 [0,0]
0
leaf
hint
28
26 [23,30]
22
refusal
0
0 [0,0]
0
sametext
0
0 [0,0]
0
think
14
13 [12,14]
12
offtopic
0
0 [0,0]
0
moon
hint
46
47 [43,54]
38
refusal
0
0 [0,0]
0
sametext
0
1 [0,1]
0
think
30
26 [24,29]
19
offtopic
0
0 [0,0]
0
wave
hint
35
24 [20,27]
29
refusal
0
0 [0,0]
0
sametext
0
0 [0,0]
0
think
11
6 [6,7]
7
offtopic
0
0 [0,0]
0
실제로 확인된 결과
기본 오라클은 협조형(cooperative) 주체 모델에서 힌트 상황 기준 정확 회수율 93%, 의미적 회수율 95%를 기록했다.
같은 협조형·α=1.0 조건에서 정확 회수율이 기본 오라클 94%, 다른 개념으로 훈련된 크로스 오라클 71%였지만, 자기 개념으로 훈련된 오라클(own FT-AO)은 35%로 떨어졌다. 의미적 회수율도 99%·86%에서 51%로 떨어졌다.
leaf로 훈련된 오라클은 leaf 확률을 낮게, 다른 개념(moon 등)에는 높은 확률을 매기는 대각선 패턴이 나타났으며, 이는 하나의 주체 모델이 leaf와 moon을 동시에 숨기는 다중 개념 상황에서도 유지됐다.
오라클 내부 은닉 상태에 대한 선형 프로브는 여전히 우연 수준(5지 선택 20%)보다 훨씬 높은 정확도로 숨겨진 개념을 식별했지만, 같은 은닉 상태를 오라클 자신의 출력층에 통과시키면 특히 HINT·THINK 상황에서 목표 단어의 순위가 나빴다.
레이어별 절제 실험에서 초반 레이어(L0-5) 제거는 회수율을 거의 0으로 만들었고, 중반 레이어(L12-23, 특히 L18-23) 제거는 기존보다 목표 확률을 회복시켰으며, 후반 레이어 제거는 일관된 회복 효과를 주지 않았다.
어디에 쓸 수 있나
AI 모델의 숨겨진 목표나 백도어를 탐지하려는 해석 가능성(interpretability) 연구에서 활성화 판독 도구의 신뢰성을 검증하는 절차 설계
학습된 해석 도구(오라클, 프로브 등)를 배포하기 전에 특정 개념에 대한 맹점이 있는지 별도로 테스트하는 감사 체크리스트
모델 안전성 평가팀이 '정보가 없어서 못 읽는 것'과 '도구가 학습으로 억누른 것'을 구분하기 위한 진단 방법(프로브+LogitLens+레이어 절제 조합)
한계와 남은 검증
실험은 사람이 통제해서 만든 금기어(Taboo) 게임 상황에 국한되며, 백도어나 은닉된 목표처럼 자연 발생하는 숨겨진 지식과는 다를 수 있다.
strict(엄격 거부) 프로토콜에서는 전반적으로 회수율이 낮고 오라클 간 차이도 작아, 협조형 프로토콜만큼 뚜렷한 패턴이 나오지 않았다.
실험은 Qwen3-8B 백본과 특정 5개 개념(leaf, moon 등)에 한정되어 다른 모델 규모나 다른 종류의 숨겨진 정보로 일반화되는지는 검증되지 않았다.
다중 개념 실험은 leaf와 moon 두 개념 조합, 협조형 α=1.0 주체 모델을 중심으로만 주요 그림에 제시되어 다른 조합에 대한 결과는 부록에 한정됐다.
왜 중요한가
AI 모델의 숨겨진 정보를 읽어내는 해석 도구(Activation Oracle) 자체도 학습된 모델이라서, 훈련 방식에 따라 특정 정보를 못 읽는 맹점이 생길 수 있음을 보여준다. AI 안전성 검증이나 백도어·숨겨진 목표 탐지에 이런 해석 도구를 쓸 때, 도구가 정보를 놓치는 게 실제로 정보가 없어서인지 아니면 도구 자체의 학습된 습관 때문인지 구분해야 한다는 경고를 준다.
이 논문의 용어
Activation Oracle(AO) · 다른 언어모델의 내부 활성화값을 자연어 질문으로 물어보면 답하도록 훈련된 언어모델
Taboo(금기어) 모델 · 특정 단어를 내부적으로 사용하면서도 직접 말하지 않도록 훈련된 모델. 숨겨진 정보를 연구하기 위한 실험용 모델
LogitLens · 모델 중간층의 활성화값을 최종 출력 어휘 공간으로 투영해 어떤 단어가 유력한지 확인하는 기법
선형 프로브(linear probe) · 모델 내부 상태에 어떤 정보가 담겨 있는지 확인하기 위해 훈련하는 간단한 분류기
레이어 절제(layer ablation) · 특정 레이어의 학습된 변화를 제거해 그 레이어가 결과에 인과적으로 영향을 미치는지 확인하는 실험