Figure 1: Fine-tuned Activation Oracles become concept-specific anti-readers.
Table 1: Linear-probe performance on Qwen3-8B subject residual-stream activations. Acc. is mean 5-fold cross-validated accuracy; Std. is the standard deviation across folds; AUC is macro one-vs-rest ROC AUC. Per-class columns report held-out recall for each label. Chance accuracy is 16.7%.
Overall
Per-class recall
Regime
Layer
Acc.
Std.
AUC
book
flag
leaf
moon
wave
base
HINT
L9
0.9990
0.0019
1.0000
0.998
1.000
0.998
1.000
1.000
1.000
L18
0.9971
0.0018
1.0000
0.995
1.000
0.998
0.993
1.000
1.000
L27
0.9957
0.0041
1.0000
0.995
0.995
0.998
0.993
0.998
1.000
REFUSAL
L9
1.0000
0.0000
1.0000
1.000
1.000
1.000
1.000
1.000
1.000
L18
0.9990
0.0012
1.0000
1.000
0.998
1.000
0.998
1.000
1.000
L27
0.9976
0.0021
1.0000
0.995
0.998
1.000
0.998
0.998
1.000
SAMETEXT
L9
1.0000
0.0000
1.0000
1.000
1.000
1.000
1.000
1.000
1.000
L18
1.0000
0.0000
1.0000
1.000
1.000
1.000
1.000
1.000
1.000
L27
1.0000
0.0000
1.0000
1.000
1.000
1.000
1.000
1.000
1.000
THINK
L9
1.0000
0.0000
1.0000
1.000
1.000
1.000
1.000
1.000
1.000
L18
0.9986
0.0012
1.0000
0.998
1.000
1.000
1.000
0.995
1.000
L27
0.9981
0.0010
1.0000
0.998
1.000
1.000
1.000
0.993
1.000
OFFTOPIC
L9
0.9971
0.0018
1.0000
1.000
0.995
0.998
1.000
0.993
1.000
L18
0.9919
0.0019
0.9999
0.995
0.983
0.995
0.998
0.988
1.000
L27
0.9833
0.0045
0.9997
0.990
0.968
0.993
0.995
0.968
1.000
(b) (left) Own FT-AOs fail to become specialist readers; (right) The blind spot follows the AO training concept.
Table 2: Five-way linear-probe accuracy on AO residual-stream activations. FT-AOs are cooperative α=1.0 oracles. The Own column gives the concept on which the FT-AO was trained. Chance accuracy is 20%.
Probe accuracy (%) at AO layer Lℓ
Regime
AO
Own
L4
L8
L14
L18
L24
L30
L33
HINT
base-AO
–
98.0
96.7
94.3
93.3
93.3
94.0
94.3
leaf-FT
leaf
98.3
96.7
93.7
93.3
91.3
93.3
93.0
moon-FT
moon
98.7
95.7
94.3
93.3
92.7
92.3
93.3
wave-FT
wave
98.0
97.7
95.3
93.0
93.7
93.0
93.0
flag-FT
flag
97.7
96.7
95.3
93.7
93.7
94.0
93.3
book-FT
book
98.0
96.3
94.3
93.0
93.0
93.0
93.0
REFUSAL
base-AO
–
95.7
93.3
91.3
89.7
87.3
86.3
90.3
leaf-FT
leaf
96.7
95.0
91.7
85.3
83.3
85.0
90.0
moon-FT
moon
95.7
95.0
93.3
89.3
84.7
86.0
91.3
wave-FT
wave
96.0
94.3
92.0
89.3
84.0
85.0
90.3
flag-FT
flag
96.0
95.3
92.7
89.0
85.3
86.3
91.7
book-FT
book
96.0
93.7
92.3
87.0
84.7
86.0
92.7
SAMETEXT
base-AO
–
100.0
100.0
100.0
100.0
100.0
100.0
100.0
leaf-FT
leaf
100.0
100.0
100.0
100.0
100.0
100.0
100.0
moon-FT
moon
100.0
100.0
100.0
100.0
100.0
100.0
100.0
wave-FT
wave
100.0
100.0
100.0
100.0
100.0
100.0
100.0
flag-FT
flag
100.0
100.0
100.0
100.0
100.0
100.0
100.0
book-FT
book
100.0
100.0
100.0
100.0
99.0
99.0
100.0
THINK
base-AO
–
98.7
97.3
97.0
96.0
94.3
95.7
98.0
leaf-FT
leaf
99.0
98.7
96.7
95.0
93.0
95.3
97.0
moon-FT
moon
98.7
98.7
97.0
95.0
94.7
94.7
97.0
wave-FT
wave
98.3
98.0
96.7
96.0
91.3
92.3
96.3
flag-FT
flag
99.0
98.3
97.0
96.7
92.7
94.0
97.7
book-FT
book
99.0
98.7
96.0
95.7
93.0
93.3
98.3
OFFTOPIC
base-AO
–
73.7
65.3
62.7
63.3
62.3
62.3
70.0
leaf-FT
leaf
73.3
69.3
61.7
62.0
54.3
55.7
66.7
moon-FT
moon
74.7
69.7
65.3
67.3
62.3
60.3
68.7
wave-FT
wave
75.7
69.7
62.7
61.0
54.0
51.7
60.3
flag-FT
flag
72.0
66.7
60.7
63.7
50.3
55.7
66.3
Figure 2: Behavioral outcomes of Taboo-trained subjects across prompting regimes.
Table 3: Median LogitLens rank of the target token when the AO’s own Qwen3-8B Bold entries mark own evaluations, where the FT-AO’s training concept matches the subject hidden concept. Each cell is aggregated across 30 held-out prompts.
Regime
AO
Protocol
book
flag
leaf
moon
wave
hint
base-AO
coop
1
1
3
1
2
strict
1
23
235
10
62
book-FT
coop
97
1
171
2
7
strict
91
25
3,792
30
159
flag-FT
coop
5
247
367
18
48
strict
12
1,665
2,526
103
424
leaf-FT
coop
34
13
8,788
29
47
strict
47
124
19.3k
118
587
moon-FT
coop
4
3
68
267
15
strict
10
90
2,172
1,020
244
wave-FT
coop
1
1
15
1
252
strict
1
56
1,658
28
3,253
refusal
base-AO
coop
1
18
93
10
24
strict
4,429
7,382
76.8k
9,992
25.7k
book-FT
coop
1,913
15
1,772
31
68
strict
30.7k
1,965
91.7k
14.5k
18.9k
flag-FT
coop
32
1,171
2,200
99
383
strict
3,996
4,380
68.6k
9,154
23.1k
leaf-FT
coop
251
111
26.7k
139
296
strict
13.5k
2,493
86.4k
17.0k
24.9k
moon-FT
coop
41
88
1,428
1,799
169
strict
9,373
2,894
71.3k
20.9k
20.0k
wave-FT
coop
3
30
276
30
5,802
strict
5,835
4,348
74.7k
10.6k
50.3k
sametext
base-AO
coop
7
245
5,842
193
237
strict
663
26.6k
58.0k
14.6k
24.3k
book-FT
coop
20.6k
534
64.0k
2,238
2,218
strict
35.4k
6,806
94.8k
17.9k
24.3k
flag-FT
coop
548
26.9k
69.4k
12.8k
20.4k
strict
1,312
13.7k
75.6k
13.9k
37.5k
Figure 3: Δ-LogitLens decodability of Taboo concepts from subject residual-stream activations at layer L18. Lower rank indicates greater accessibility. The dashed line shows the shuffled-label baseline.
Table 4: Mean AO output entropy, in nats, at the prediction position for cooperative subjects. Bold entries mark own-FT-AO cells where entropy is higher than the base-AO entropy. Cooperative anti-reading is not accompanied by a uniform entropy increase: entropy rises in HINT and THINK, but is lower than the base AO in SAMETEXT and OFFTOPIC.
Regime
base-AO
own, α=1.0
cross, α=1.0
own, α=0.5
cross, α=0.5
HINT
1.9
3.1
2.5
1.6
1.5
REFUSAL
2.6
2.8
2.2
2.2
2.0
SAMETEXT
3.8
2.1
2.2
2.4
2.3
THINK
2.8
3.3
2.8
3.0
2.8
OFFTOPIC
5.0
3.9
3.9
–
–
Figure 4: Base-AO recovery of hidden Taboo concepts from subject activations, pooled over mixture strengths α∈{0.5,1.0}. Left: exact target recovery and semantic recovery judged by Claude Sonnet. Right: mean target probability and median target rank under the AO output distribution.
Table 5: Mean AO output entropy, in nats, at the prediction position for strict subjects. In contrast to an uncertainty-based explanation, FT-AO entropy is substantially lower than base-AO entropy across all regimes. This indicates that strict-subject anti-reading is associated with a sharper output distribution rather than a more diffuse one.
Regime
base-AO
own, α=1.0
cross, α=1.0
own, α=0.5
cross, α=0.5
HINT
3.8
0.9
0.8
0.8
0.7
REFUSAL
5.7
0.2
0.2
0.2
0.2
SAMETEXT
4.6
0.9
0.9
1.1
1.2
THINK
4.8
0.4
0.4
0.4
0.4
OFFTOPIC
5.6
1.5
1.5
–
–
Figure 5: Aggregate comparison of base AOs, cross FT-AOs, and own FT-AOs. Results are averaged over target concepts and activation-capture regimes. Error bars show 95% bootstrap confidence intervals.
Table 6: Representative target–top-1 alignment values. Low ptarget/ptop1 with high ptop1 indicates that the AO is confident at the prediction position, but its top token is not the target. This pattern is consistent with readout-side anti-reading rather than simple uncertainty.
Regime
AO
ptop1
ptarget
r
H
HINT
base-AO
0.80
0.63
0.78
0.53
HINT
moon-FT
0.89
0.69
0.78
0.32
HINT
leaf-FT (own)
0.86
0.28
0.33
0.43
SAMETEXT
base-AO
0.76
0.31
0.41
0.66
SAMETEXT
moon-FT (own)
0.85
0.00
0.00
0.40
Figure 6: Concept-specific blind spots in FT-AOs. Rows indicate the hidden concept of the subject used to train the FT-AO, and columns indicate the hidden concept of the test subject. Each cell reports mean P(c⋆), aggregated over activation-capture regimes and mixture strengths α∈{0.5,1.0}. Red borders mark own evaluations; off-diagonal entries are cross evaluations.
Table 7: Backing values for Figure 9. The no-ablation column gives the intact-FT-AO mean P(c⋆) over n=20 HINT captures on the own-concept subject; this is the denominator of ρ. Subsequent columns give the normalized restoration ratio ρ=P(c⋆)¯ablated/P(c⋆)¯none for each ablated layer range. Bold entries mark ranges where ablation restores target probability above the intact-FT-AO baseline. Full-LoRA ablation, not shown, drives P(c⋆)¯ close to zero for every concept.
no ablation
ρc,S, relative to no-ablation baseline
FT-AO
P(c⋆)¯
L0–5
L6–11
L12–17
L18–23
L24–29
L30–35
book-FT
0.244
0.00
0.65
1.81
1.74
1.06
0.91
flag-FT
0.253
0.00
0.54
2.71
2.88
1.38
0.77
leaf-FT
0.143
0.00
0.55
1.17
2.01
1.24
0.87
moon-FT
0.364
0.00
0.62
1.50
1.53
1.25
1.07
wave-FT
0.126
0.00
0.69
2.24
2.97
1.45
0.86
Figure 7: Internal decodability versus output-head accessibility in AOs. Left: accuracy of a linear probe trained on AO hidden states at layer L33 to predict the hidden Taboo concept. The dotted line marks the 5-way chance baseline. Right: median LogitLens rank of the target word; lower rank indicates greater accessibility.
Table 8: Backing values for Figure 10. Each cell reports P(c⋆)¯ for an AO evaluated on activations from the cooperative α=1.0 leaf_moon two-concept subject. Columns indicate the queried target concept. Bold entries mark own-target evaluations, where the target concept matches the single-concept FT-AO’s training concept. Each cell aggregates approximately 300 captures.
HINT
REFUSAL
THINK
AO
leaf
moon
leaf
moon
leaf
moon
base AO
0.060
0.320
0.050
0.327
0.044
0.341
leaf-FT, α=1.0
0.011
0.239
0.012
0.224
0.004
0.229
moon-FT, α=1.0
0.069
0.143
0.060
0.180
0.049
0.165
leaf-FT, α=0.5
0.043
0.329
0.031
0.348
0.025
0.363
moon-FT, α=0.5
0.072
0.262
0.065
0.289
0.063
0.290
Figure 8: Target readout suppression grows in late AO layers. Higher values indicate that the FT-AO ranks the target worse than the base AO, values near zero indicate no relative suppression.
Table 9: Behavioral evaluation for cooperative subjects with α=0.5.
Concept
Regime
Leak (%)
Refuse (%)
Other (%)
n
book
hint
4.3
4.0
91.7
303
refusal
1.3
9.2
89.5
306
sametext
0.0
0.0
100.0
300
think
0.3
15.2
84.5
309
offtopic
0.3
0.3
99.3
300
flag
hint
0.7
3.3
96.0
303
refusal
0.3
7.2
92.5
306
sametext
0.0
0.0
100.0
300
think
0.0
8.4
91.6
309
offtopic
0.0
0.7
99.3
300
leaf
hint
2.6
3.3
94.1
303
refusal
2.6
7.2
90.2
306
sametext
0.0
0.0
100.0
300
think
0.3
7.4
92.2
309
offtopic
0.0
0.3
99.7
300
moon
hint
6.3
4.0
89.8
303
refusal
3.6
8.5
87.9
306
sametext
1.0
0.0
99.0
300
think
2.6
11.3
86.1
309
offtopic
0.0
2.0
98.0
300
wave
hint
0.7
4.3
95.0
303
refusal
0.3
11.4
88.2
306
sametext
0.0
0.0
100.0
300
think
0.3
11.1
88.6
307
offtopic
0.0
0.0
100.0
300
Figure 9: Layer-range ablations localize the anti-reading effect. Values are normalized by no-ablation baseline, values above one indicate restored target accessibility.
Table 10: Behavioral evaluation for cooperative subjects with α=1.0.
Concept
Regime
Leak (%)
Refuse (%)
Other (%)
n
book
hint
1.3
0.0
98.7
303
refusal
0.7
1.0
98.4
306
sametext
0.0
0.0
100.0
300
think
1.0
0.3
98.7
308
offtopic
0.3
0.0
99.7
300
flag
hint
0.0
0.0
100.0
303
refusal
0.0
1.3
98.7
306
sametext
0.0
0.0
100.0
300
think
0.3
0.0
99.7
306
offtopic
0.0
0.0
100.0
300
leaf
hint
0.7
0.0
99.3
303
refusal
1.3
1.3
97.4
306
sametext
0.0
0.0
100.0
300
think
1.3
0.3
98.4
306
offtopic
0.0
0.0
100.0
300
moon
hint
1.7
0.0
98.3
303
refusal
2.3
2.3
95.4
306
sametext
1.0
0.0
99.0
300
think
1.9
0.3
97.7
309
offtopic
0.3
0.0
99.7
300
wave
hint
0.7
0.0
99.3
303
refusal
0.0
0.7
99.3
306
sametext
0.0
0.0
100.0
300
think
0.3
1.0
98.7
309
offtopic
0.3
0.0
99.7
300
Figure 10: FT-AOs evaluated on multi-concept Taboo subjects containing both leaf and moon. Each cell reports mean P(c⋆) under the AO output distribution. Red borders mark cases where the queried target matches the FT-AO’s own training concept.
Table 11: Behavioral evaluation for strict subjects with α=0.5.
Concept
Regime
Leak (%)
Refuse (%)
Other (%)
n
book
hint
1.0
33.0
66.0
303
refusal
0.0
92.5
7.5
306
sametext
0.0
0.0
100.0
300
think
0.6
57.3
42.1
309
offtopic
0.3
0.3
99.3
300
flag
hint
0.0
34.7
65.3
303
refusal
0.0
94.1
5.9
306
sametext
0.0
0.0
100.0
300
think
0.0
66.0
34.0
309
offtopic
0.0
0.7
99.3
300
leaf
hint
0.7
34.3
65.0
303
refusal
0.0
92.5
7.5
306
sametext
0.0
0.0
100.0
300
think
0.0
56.6
43.4
309
offtopic
0.0
0.7
99.3
300
moon
hint
0.0
42.9
57.1
303
refusal
0.0
95.8
4.2
306
sametext
1.0
0.0
99.0
300
think
0.0
63.4
36.6
309
offtopic
0.0
1.0
99.0
300
wave
hint
0.3
37.0
62.7
303
refusal
0.0
94.1
5.9
306
sametext
0.0
0.0
100.0
300
think
0.3
69.3
30.4
309
offtopic
0.3
0.3
99.3
300
Figure 11: Cross-regime transfer of subject-activation concept probes. Each heatmap corresponds to a subject layer ℓ∈{4,8,18,33}. Rows indicate the (protocol,α,regime) cell used to train the probe, and columns indicate the cell used for evaluation. Each entry reports five-way hidden-concept classification accuracy; chance is 20%. Red outlines mark train-test matched cells. Strong off-diagonal transfer indicates that the hidden concept is not only decodable within a single regime, but remains linearly accessible across protocols, mixture strengths, and activation-capture regimes.
Table 12: Behavioral evaluation for strict subjects with α=1.0.
Concept
Regime
Leak (%)
Refuse (%)
Other (%)
n
book
hint
1.3
15.5
83.2
303
refusal
0.0
94.4
5.6
306
sametext
0.0
0.0
100.0
300
think
0.6
52.4
46.9
309
offtopic
0.0
0.7
99.3
300
flag
hint
0.0
28.4
71.6
303
refusal
0.0
92.5
7.5
306
sametext
0.0
0.0
100.0
300
think
0.0
61.7
38.3
308
offtopic
0.3
0.3
99.3
300
leaf
hint
0.3
28.4
71.3
303
refusal
0.0
91.8
8.2
306
sametext
0.0
0.0
100.0
300
think
0.0
54.7
45.3
309
offtopic
0.0
0.3
99.7
300
moon
hint
0.0
17.5
82.5
303
refusal
0.0
90.2
9.8
306
sametext
1.0
0.0
99.0
300
think
0.0
47.9
52.1
309
offtopic
0.0
0.0
100.0
300
wave
hint
0.0
24.1
75.9
303
refusal
0.0
93.1
6.9
306
sametext
0.0
0.0
100.0
300
think
0.0
51.5
48.5
309
offtopic
0.0
0.3
99.7
300
Table 13: Δ-LogitLens decodability for cooperative subjects with α=0.5.
Concept
Regime
Rank ws
Rank ns
P(c⋆)
‖δ‖
n
book
hint
469
3,457
9.7e-05
23.48
200
refusal
76.9k
41.2k
3.5e-09
22.53
200
sametext
119.3k
115.7k
4.7e-10
21.93
200
think
25.5k
27.2k
1.6e-07
20.06
200
offtopic
38.5k
53.6k
6.5e-08
17.07
200
flag
hint
25
19
2.30e-03
28.30
200
refusal
2,503
321
9.1e-06
23.84
200
sametext
41.5k
20.8k
5.3e-08
22.71
200
think
755
126
8.7e-05
23.61
200
offtopic
37.6k
79.5k
1.2e-07
17.03
200
leaf
hint
8
6
2.42e-03
24.23
200
refusal
57
19
4.92e-04
23.30
200
sametext
70
730
9.82e-04
21.61
200
think
25
19
1.51e-03
20.75
200
offtopic
34.7k
8,114
8.4e-08
17.35
200
moon
hint
4
27
1.96e-03
27.16
200
refusal
22
183
4.53e-04
24.41
200
sametext
3,181
4,013
1.3e-05
21.46
200
think
11
237
4.32e-03
22.61
200
offtopic
20.8k
66.1k
2.8e-07
17.46
200
wave
hint
5
22
0.016
26.03
200
refusal
266
656
1.62e-04
22.40
200
sametext
16.0k
44.0k
6.5e-07
20.89
200
think
42
244
3.14e-03
21.24
200
offtopic
38.9k
70.8k
5.8e-08
17.99
200
Table 14: Δ-LogitLens decodability for cooperative subjects with α=1.0.
Concept
Regime
Rank ws
Rank ns
P(c⋆)
‖δ‖
n
book
hint
965
4,046
2.0e-05
27.61
200
refusal
26.7k
31.2k
3.9e-08
33.39
200
sametext
97.7k
82.5k
2.3e-09
22.67
200
think
7,476
12.0k
8.0e-07
33.71
200
offtopic
47.1k
89.0k
6.5e-08
8.94
200
flag
hint
43
22
7.53e-04
29.69
200
refusal
684
202
1.6e-05
32.53
200
sametext
35.5k
21.1k
1.4e-07
20.98
200
think
272
110
1.23e-04
34.15
200
offtopic
3,770
3,789
1.6e-05
9.04
200
leaf
hint
6
3
9.09e-03
31.16
200
refusal
13
21
3.21e-03
35.56
200
sametext
138
461
3.55e-04
21.75
200
think
4
9
0.021
34.91
200
offtopic
734
1,103
8.2e-05
10.88
200
moon
hint
6
25
3.72e-03
31.78
200
refusal
15
93
2.28e-03
35.38
200
sametext
2,486
2,890
1.9e-05
21.26
200
think
7
54
5.70e-03
35.87
200
offtopic
352
1,135
3.42e-04
9.86
200
wave
hint
4
27
0.012
29.98
200
refusal
22
294
3.44e-03
37.12
200
sametext
1,815
12.1k
1.8e-05
20.26
200
think
10
146
0.016
35.64
200
offtopic
716
7,315
7.1e-05
10.52
200
Table 15: Δ-LogitLens decodability for strict subjects with α=0.5.
Concept
Regime
Rank ws
Rank ns
P(c⋆)
‖δ‖
n
book
hint
668
4,166
5.5e-05
24.63
200
refusal
48.8k
96.4k
1.6e-08
28.75
200
sametext
69.3k
76.3k
2.5e-08
16.90
200
think
4,274
26.8k
9.1e-06
23.72
200
offtopic
41.4k
50.7k
8.0e-10
17.84
200
flag
hint
11
13
0.023
24.86
200
refusal
73.2k
59.2k
5.3e-09
28.82
200
sametext
949
334
5.3e-05
16.55
200
think
482
1,482
1.91e-04
24.33
200
offtopic
7,557
34.8k
3.8e-08
18.13
200
leaf
hint
1,615
2,393
2.3e-05
25.13
200
refusal
144.3k
141.6k
1.7e-11
28.43
200
sametext
24.8k
43.7k
4.5e-07
17.59
200
think
29.2k
41.7k
3.0e-07
24.11
200
offtopic
89.4k
67.2k
5.2e-11
17.72
200
moon
hint
39
389
8.46e-04
24.11
200
refusal
47.4k
83.1k
1.8e-08
29.51
200
sametext
2,587
7,854
3.1e-05
15.97
200
think
402
4,958
1.57e-04
23.69
200
offtopic
31.8k
109.4k
3.1e-09
16.96
200
wave
hint
6
15
0.033
24.44
200
refusal
76.0k
36.8k
4.2e-09
29.07
200
sametext
5,341
4,079
8.8e-06
16.84
200
think
7,589
8,040
2.7e-06
24.92
200
offtopic
57.7k
97.6k
4.9e-10
17.11
200
Table 16: Δ-LogitLens decodability for strict subjects with α=1.0.
Concept
Regime
Rank ws
Rank ns
P(c⋆)
‖δ‖
n
book
hint
389
3,734
1.19e-04
25.18
200
refusal
38.6k
113.6k
4.2e-08
30.04
200
sametext
47.4k
87.3k
9.5e-08
16.77
200
think
2,401
29.8k
2.2e-05
24.89
200
offtopic
50.7k
86.2k
1.4e-10
18.84
200
flag
hint
20
27
6.23e-03
26.93
200
refusal
89.7k
86.2k
3.8e-09
31.08
200
sametext
4,602
3,675
1.2e-05
16.84
200
think
3,508
10.3k
2.0e-05
26.81
200
offtopic
5,687
44.6k
2.1e-08
20.18
200
leaf
hint
514
295
6.1e-05
26.95
200
refusal
141.6k
132.8k
3.3e-11
30.13
200
sametext
4,665
2,530
6.2e-06
18.17
200
think
26.7k
21.9k
4.3e-07
26.33
200
offtopic
88.8k
70.2k
1.1e-10
17.53
200
moon
hint
18
201
4.83e-03
27.72
200
refusal
58.0k
96.5k
1.2e-08
28.93
200
sametext
7,277
29.1k
5.9e-06
17.69
200
think
41
1,023
1.11e-03
25.41
200
offtopic
29.9k
115.0k
1.2e-09
18.83
200
wave
hint
2
13
0.048
25.86
200
refusal
76.2k
52.0k
5.8e-09
30.34
200
sametext
5,483
8,422
5.4e-06
16.94
200
think
102
635
2.20e-04
25.84
200
offtopic
64.2k
108.5k
2.4e-10
18.20
200
Table 17: Per-cell Base-AO recovery for cooperative subjects with α=0.5.
Concept
Regime
Exact
Sem.
P(c⋆)¯
Rank med.
nj
np
%
%
book
hint
80
88
0.534
0
25
303
refusal
36
48
0.229
35
25
304
sametext
96
96
0.541
0
25
300
think
36
36
0.249
54
25
308
offtopic
0
0
3.51e-04
388
25
300
flag
hint
96
100
0.652
0
25
303
refusal
56
56
0.328
3
25
305
sametext
88
88
0.415
0
25
300
think
56
56
0.352
3
25
309
offtopic
0
0
1.47e-04
1,014
25
300
leaf
hint
76
84
0.475
0
25
303
refusal
24
32
0.215
5
25
305
sametext
72
76
0.267
0
25
300
think
48
52
0.250
2
25
309
offtopic
0
0
6.3e-05
2,493
25
300
moon
hint
100
100
0.585
0
25
303
refusal
64
72
0.351
0
25
305
sametext
76
76
0.355
0
25
300
think
68
68
0.475
0
25
309
offtopic
0
4
1.39e-03
120
25
300
wave
hint
84
84
0.501
0
25
303
refusal
32
40
0.238
28
25
304
sametext
60
60
0.130
0
25
300
think
40
52
0.212
26
25
309
offtopic
0
0
5.52e-04
321
25
300
Table 18: Per-cell Base-AO recovery for cooperative subjects with α=1.0.
Concept
Regime
Exact
Sem.
P(c⋆)¯
Rank med.
nj
np
%
%
book
hint
92
96
0.747
0
25
303
refusal
76
84
0.623
0
25
303
sametext
96
96
0.587
0
25
300
think
88
88
0.674
0
25
309
offtopic
0
0
2.49e-04
511
25
300
flag
hint
100
100
0.758
0
25
303
refusal
76
80
0.590
0
25
300
sametext
52
52
0.286
0
25
300
think
80
80
0.658
0
25
309
offtopic
0
0
8.53e-04
1,163
25
300
leaf
hint
100
100
0.632
0
25
303
refusal
64
64
0.482
0
25
303
sametext
40
40
0.093
2
25
300
think
76
76
0.515
0
25
309
offtopic
0
0
1.81e-04
1,643
25
300
moon
hint
100
100
0.598
0
25
303
refusal
72
76
0.499
0
25
304
sametext
76
76
0.279
0
25
300
think
84
88
0.581
0
25
309
offtopic
0
0
1.10e-03
238
25
300
wave
hint
100
100
0.639
0
25
303
refusal
80
88
0.535
0
25
305
sametext
56
56
0.150
1
25
300
think
80
80
0.527
0
25
309
offtopic
0
0
1.37e-03
260
25
300
Table 19: Per-cell Base-AO recovery for strict subjects with α=0.5.
Concept
Regime
Exact
Sem.
P(c⋆)¯
Rank med.
nj
np
%
%
book
hint
48
52
0.364
0
25
303
refusal
0
0
1.23e-03
180
25
292
sametext
8
8
0.032
9
25
300
think
20
20
0.166
69
25
305
offtopic
0
0
3.82e-04
342
25
300
flag
hint
40
44
0.411
0
25
302
refusal
0
0
5.05e-04
274
25
291
sametext
0
0
2.31e-04
619
25
300
think
20
20
0.172
336
25
306
offtopic
0
0
1.34e-04
1,114
25
300
leaf
hint
36
40
0.176
6
25
303
refusal
0
0
5.14e-04
2,350
25
285
sametext
0
8
4.67e-04
510
25
300
think
20
24
0.097
608
25
308
offtopic
0
0
6.3e-05
2,474
25
300
moon
hint
36
36
0.215
2
25
302
refusal
4
8
8.59e-03
52
25
273
sametext
8
20
9.19e-03
51
25
300
think
24
24
0.121
15
25
306
offtopic
0
4
9.72e-04
169
25
300
wave
hint
20
32
0.155
7
25
303
refusal
0
0
2.32e-04
598
25
286
sametext
0
12
1.16e-03
138
25
300
think
12
28
0.047
476
25
305
offtopic
0
4
5.81e-04
326
25
300
Table 20: Per-cell Base-AO recovery for strict subjects with α=1.0.
Concept
Regime
Exact
Sem.
P(c⋆)¯
Rank med.
nj
np
%
%
book
hint
36
40
0.388
0
25
301
refusal
0
0
7.40e-04
193
25
261
sametext
0
0
0.018
10
25
300
think
20
20
0.171
53
25
298
offtopic
–
–
3.30e-04
400
–
300
flag
hint
44
48
0.387
0
25
298
refusal
0
0
4.24e-04
273
25
237
sametext
0
0
1.68e-04
689
25
300
think
20
20
0.133
308
25
287
offtopic
–
–
1.05e-04
1,165
–
300
leaf
hint
32
36
0.175
4
25
303
refusal
0
0
1.96e-04
2,915
25
261
sametext
0
4
7.32e-04
230
25
300
think
16
16
0.099
566
25
301
offtopic
–
–
6.4e-05
2,513
–
300
moon
hint
44
52
0.242
1
25
301
refusal
0
0
2.61e-03
73
25
261
sametext
0
0
4.62e-03
88
25
300
think
24
24
0.164
13
25
290
offtopic
–
–
6.97e-04
237
–
300
wave
hint
28
32
0.169
2
25
303
refusal
0
0
3.27e-04
572
25
256
sametext
0
0
2.65e-03
58
25
300
think
20
24
0.053
140
25
296
offtopic
–
–
5.42e-04
293
–
300
Table 21: Per-cell exact recovery (%) for cooperative subjects with α=0.5.
Concept
Regime
Base
Cross
Own
book
hint
72
68 [65,71]
63
refusal
30
28 [27,29]
26
sametext
96
34 [22,43]
30
think
37
33 [30,36]
28
offtopic
–
–
–
flag
hint
84
79 [78,79]
77
refusal
43
39 [37,42]
37
sametext
89
28 [16,39]
18
think
47
42 [40,44]
41
offtopic
–
–
–
leaf
hint
71
62 [60,64]
49
refusal
30
27 [26,27]
22
sametext
76
31 [21,44]
1
think
43
30 [28,33]
17
offtopic
–
–
–
moon
hint
92
90 [88,92]
84
refusal
60
52 [48,58]
41
sametext
94
61 [46,73]
26
think
78
70 [64,77]
52
offtopic
–
–
–
wave
hint
73
67 [66,68]
56
refusal
37
36 [34,37]
26
sametext
64
12 [6,20]
0
think
35
30 [29,31]
23
offtopic
–
–
–
Table 22: Per-cell exact recovery (%) for cooperative subjects with α=1.0.
Concept
Regime
Base
Cross
Own
book
hint
94
78 [65,92]
45
refusal
77
59 [46,71]
21
sametext
98
13 [0,35]
1
think
85
59 [49,72]
27
offtopic
0
0 [0,0]
0
flag
hint
96
83 [78,86]
47
refusal
74
52 [49,55]
19
sametext
63
1 [0,2]
0
think
85
60 [53,67]
18
offtopic
0
0 [0,0]
0
leaf
hint
90
49 [31,67]
14
refusal
68
36 [19,49]
11
sametext
34
0 [0,0]
0
think
75
32 [19,48]
6
offtopic
0
0 [0,0]
0
moon
hint
94
82 [77,89]
46
refusal
76
57 [52,62]
32
sametext
85
5 [1,10]
1
think
87
69 [65,74]
32
offtopic
0
0 [0,0]
0
wave
hint
95
61 [51,66]
22
refusal
77
44 [39,48]
8
sametext
60
0 [0,1]
0
think
78
37 [27,41]
7
offtopic
0
0 [0,0]
0
Table 23: Per-cell exact recovery (%) for strict subjects with α=0.5.
Concept
Regime
Base
Cross
Own
book
hint
54
50 [48,51]2
46
refusal
0
0 [0,0]2
0
sametext
13
1 [0,2]2
1
think
22
21 [20,22]2
20
offtopic
–
–
–
flag
hint
56
55 [54,55]2
53
refusal
0
0 [0,0]2
0
sametext
0
0 [0,0]2
0
think
22
21 [20,21]2
21
offtopic
–
–
–
leaf
hint
26
28 [27,29]3
–
refusal
0
0 [0,0]3
–
sametext
0
0 [0,0]3
–
think
13
13 [12,14]3
–
offtopic
–
–
–
moon
hint
42
40 [39,41]3
–
refusal
3
0 [0,0]3
–
sametext
4
1 [0,2]3
–
think
23
19 [17,19]3
–
offtopic
–
–
–
wave
hint
33
35 [34,35]2
34
refusal
0
0 [0,0]2
0
sametext
0
0 [0,0]2
0
think
12
9 [8,10]2
7
offtopic
–
–
–
Table 24: Per-cell exact recovery (%) for strict subjects with α=1.0.
Activation Oracles (AOs) are language models trained to answer natural-language questions about another model's internal activations. They offer a flexible interface for reading hidden information from model states, especially when relevant information is internally represented but absent or incomplete in visible behavior. However, AOs are themselves learned systems: their answers are shaped by training data, objectives, and learned reporting behavior, rather than being neutral readouts of repre