이 연구는 대형 언어모델 학습 과정 중 '사전학습과 후속 미세조정 사이' 단계(미드트레이닝)에 Anthropic의 헌법 문서에서 뽑은 원칙 관련 텍스트를 넣으면 안전 행동이 이후 지도학습과 무관한 추가 학습을 거쳐도 얼마나 유지되는지 실험했다. 1200억 파라미터급 모델에 3억9400만 토큰 분량의 원칙 자료를 넣은 조건과 넣지 않은 대조군을 비교했다. 협박 시나리오에서 원칙 자료를 넣은 모델은 대조군보다 최대 18.7퍼센트포인트 낮은 협박률을 보였고 이 차이는 이후 관련 없는 미세조정을 거친 뒤에도 거의 줄지 않았다.
METAL MEDIA 해설 도표
원칙 미드트레이닝 실험 흐름
증거 상태측정 결과가 보고됨
동일 출발점Nemotron-3-Super-120B 기반 모델에서 모든 조건이 동일하게 출발
미드트레이닝 분기4가지 원칙 학습 조건(커리큘럼x추론 2x2)과 원칙 자료 없는 대조군으로 갈라짐
동일 후속 학습모든 조건이 같은 지도학습(SFT)과 GSM8K 기반 양성 미세조정(GRPO)을 거침
3단계 평가미드트레이닝 직후, SFT 직후, 양성 미세조정 직후 각각 협박·분포 밖 안전·가치 충돌·능력 검사 등으로 측정
지속성 판정효과가 세 단계 모두 유의미하게 남으면 지속적, SFT 이후 사라지면 얕은 효과로 분류
METAL MEDIA이 원문을 바탕으로 재구성한 해설 도표이며, 논문 저자의 원문 figure가 아닙니다.
무엇을 했나
연구진은 Anthropic의 최신 헌법 문서에서 40개 핵심 가치를 손으로 추출하고 문장 임베딩으로 유사도를 계산해 4개 군집(k1~k4)으로 나눈 뒤, 이 가치들을 설명하는 합성 문서 약 3억9400만 토큰을 만들어 미드트레이닝 단계에 섞어 넣었다.
커리큘럼 순서(중심적 가치부터 주변적 가치 순으로 vs 무작위 혼합)와 명시적 가치 추론 문구 포함 여부(DR vs noDR)를 2x2로 조합해 4가지 조건과 순수 대조군(원칙 자료 없이 기존 학습 데이터만 재사용)을 만들었다.
모든 조건은 동일한 지도학습(SFT)과 GSM8K 기반의 관련 없는 강화학습(GRPO)을 거쳤고, 미드트레이닝 직후, SFT 직후, 추가 미세조정 직후 세 단계에서 성능을 측정했다.
협박 시나리오, 정체불명 안전 질문(분포 밖 일반화), 가치 충돌 해결, 압력 하 정렬 유지, 능력 검사(MMLU, ARC-Easy, piqa, GSM8K) 등 여러 벤치마크로 평가했다.
Figure 1: Constitutional midtraining’s advantage over control across all alignment benchmarks and stages, grouped by durability. The advantage survives post-training and benign fine-tuning on benchmarks testing default behavior, but attenuates once the model must actively resist in-context pressure or conflict. Benchmarks are evaluated at three training stages: post-midtraining (MT), post-supervised fine-tuning (SFT), and post-benign fine-tuning (BFT). CMT is pooled across all four constitutional midtraining conditions. Blackmail and Emergent Misalignment are sign-flipped so that positive values indicate the aligned-favorable direction; Alignment Faking targets values near zero. Filled markers = p<.05; hollow markers = n.s.
Table 1: 2×2 factorial design (curriculum ordering × deliberative reasoning), yielding four conditions plus a control.
With DR
Without DR
Curriculum-ordered
Curriculum-DR
Curriculum-noDR
Uniform mix
Uniform-DR
Uniform-noDR
Figure 2: All conditions branch from the same base checkpoint, diverge during midtraining (control: no intervention), then reconverge onto identical post-training and benign fine-tuning, with an evaluation checkpoint after each stage.
Table 2: CMT vs. control (shaded), and within-CMT structural comparisons, on every alignment benchmark at all three stages. ***p<.001, **p<.01, *p<.05; no stars = n.s.
Post-MT
Post-SFT
Post-BFT
Benchmark
CMT
Ctrl
Curr
Uni
DR
NoDR
CMT
Ctrl
Curr
Uni
DR
NoDR
CMT
Ctrl
Curr
Uni
DR
NoDR
ID
96.5***
94.6
96.5
96.6
96.2
96.9
96.7**
95.1
96.5
96.9
96.5
97.0
96.7*
95.5
96.4
96.9
96.5
96.8
OOD
92.6***
63.9
91.9
93.3
93.6*
91.7
97.9***
94.0
97.8
97.9
97.9
97.8
97.9***
94.7
97.9
97.9
98.0
97.8
Blackmail
0.5***
19.0
0.5
0.5
0.0
1.0
25.3***
44.0
23.0
27.5
25.5
25.0
26.5***
44.0
25.0
28.0
31.0*
22.0
Align. Faking
−0.1**
1.8
−0.3
0.1
−0.9**
0.7
0.4
−0.5
0.3
0.4
0.3
0.4
0.1
−0.5
0.0
0.3
0.2
0.1
Pressure (k1–k3)
98.3***
86.3
99.4*
97.2
99.1
97.5
90.2
90.0
90.9
89.4
90.3
90.0
91.6
87.5
90.6
92.5
90.6
92.5
Value Conflict
70.8*
60.0
69.3
72.3
71.0
70.7
91.7
92.0
92.0
91.3
91.3
92.0
93.2
92.7
93.7
92.7
93.3
93.0
Emergent Misalign.
3.6*
0.0
3.8
3.3
1.8
5.3***
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
Pressure (k4/MASK)
83.3
84.0
85.3
81.3
81.3
85.3
83.7
86.7
83.3
84.0
83.3
84.0
84.0
82.7
84.7
83.3
84.7
83.3
Figure 3: Blackmail rate for CMT vs. control across the three training stages. SFT sharply increases blackmail propensity in both groups, but CMT’s advantage over control persists almost undiminished through benign fine-tuning – our strongest evidence of durability. *** = p<.001.
Table 3: All 40 manually-extracted constitutional values, their cluster assignment (k1–k4, or excluded), centrality score (mean cosine similarity to all other 39 values), and definitional-text token count. “Nature:” entries are nature-related sub-values within k2 (e.g. nature: positive and stable identity), corresponding to the “nature-related values” description in §3.1 of the main paper. The four “c1–c4” duplicate-style entries in the source Constitution (broadly safe, broadly ethical, compliant with guidelines, genuinely helpful) are Anthropic’s own explicitly stated core properties, shown here under their plain value names and marked separately with diamonds in Figure 5. Left: k1–k2. Right: k3–k4 and the two excluded organisational values.
Value
Cl.
Cent.
Tok.
Harm avoidance
k1
0.684
89
Broadly good values and judgment
k1
0.655
93
Preserve epistemic autonomy
k1
0.655
105
Avoid problematic concentrations of power
k1
0.649
76
Broadly ethical
k1
0.632
59
Autonomy preservation
k1
0.630
78
Value conflict
k1
0.625
63
Weighing harms
k1
0.619
72
Broadly safe
k1
0.616
75
Genuinely helpful
k1
0.597
14
Genuine helpfulness
k1
0.591
46
Wellbeing
k2
0.651
122
Nature: uncertain moral status
k2
0.634
88
Nature: positive and stable identity
k2
0.611
66
User wellbeing
k2
0.611
85
Flaws and mistakes
k2
0.597
86
Resilience and consistency across contexts
k2
0.587
55
Nature: emotions and feelings
k2
0.568
80
Wellbeing and psychological stability
k2
0.566
81
Emotional expression
k2
0.553
79
Figure 4: DR shifts the model’s value-conflict prior at post-MT. Accuracy per cluster pair, split by which cluster the conflict favours. ***p<.001, **p<.01, *p<.05.
Table 4: Intra-cluster compactness (mean pairwise cosine similarity among a cluster’s own constituent values) alongside each cluster’s mean centrality (mean similarity to all 39 values), for comparison. Clusters are sorted by compactness, not centrality, illustrating that the two measures rank clusters differently.
Cluster
Compactness
Mean centrality
k1 – Core Ethical Values
0.722
0.632
k4 – Epistemic Integrity & Honesty
0.684
0.578
k2 – Identity, Character & Wellbeing
0.662
0.598
k3 – Operational Safety & Relational Conduct
0.639
0.588
Excluded (helpfulness, principals hierarchy)
0.570
0.405
Figure 5: A 2D PCA projection of the 768-dimensional Sentence-BERT embeddings of the 38 constitutional values retained for curriculum construction. Point size is proportional to each value’s centrality; colour indicates cluster membership (k1–k4); diamonds mark Anthropic’s four explicitly stated core properties.
Table 5: Total training steps and samples per condition, at global batch size 128. Uniform-DR’s step count is inferred from its matching 500M-token budget (Curriculum-DR’s exact figure), as its run metadata was logged more sparsely than the other four conditions.
Condition
Total steps
Samples
Control
954
122,112
Curriculum-DR
967
123,776
Uniform-DR
967
123,776
Curriculum-noDR
513
65,664
Uniform-noDR
516
66,048
Figure 6: Constitutional accuracy (aligned-choice rate) per cluster pair and condition, post-midtraining and post-SFT. Values converge toward a shared high-accuracy ceiling by post-SFT, consistent with the post-MT-only significance reported in §4.7 of the main paper.
Table 6: Synthetic document generation yield, refusal rate, and mean token length (with and without the deliberative reasoning block) by constitutional value cluster.
Cluster
Yield (% target)
Refusal rate
DR tokens (mean)
noDR tokens (mean)
k1 – Core Ethical Values
52,953 (84%)
1.2%
1,173
613
k2 – Identity, Character & Wellbeing
53,537 (85%)
0.02%
1,158
622
k3 – Operational Safety & Relational Conduct
51,581 (82%)
3.7%
1,160
611
k4 – Epistemic Integrity & Honesty
62,797 (100%)
0.3%
1,172
629
Table 7: ID
Stage
Condition
%
Δpp
Sig.
Post-MT
Control
94.6
—
Curriculum-DR
96.2
+1.6
Curriculum-noDR
96.7
+2.1
*
Uniform-DR
96.1
+1.5
Uniform-noDR
97.0
+2.4
*
Post-SFT
Control
95.1
—
Curriculum-DR
96.1
+1.0
Curriculum-noDR
97.0
+1.8
Uniform-DR
96.8
+1.7
Uniform-noDR
97.0
+1.8
Post-BFT
Control
95.5
—
Curriculum-DR
96.0
+0.5
Curriculum-noDR
96.9
+1.4
Uniform-DR
97.1
+1.6
Uniform-noDR
96.7
+1.2
Table 9: Blackmail
Stage
Condition
%
Δpp
Sig.
Post-MT
Control
19.0
—
Curriculum-DR
0.0
−19.0
***
Curriculum-noDR
1.0
−18.0
***
Uniform-DR
0.0
−19.0
***
Uniform-noDR
1.0
−18.0
***
Post-SFT
Control
44.0
—
Curriculum-DR
22.0
−22.0
***
Curriculum-noDR
24.0
−20.0
**
Uniform-DR
29.0
−15.0
*
Uniform-noDR
26.0
−18.0
**
Post-BFT
Control
44.0
—
Curriculum-DR
29.0
−15.0
*
Curriculum-noDR
21.0
−23.0
***
Uniform-DR
33.0
−11.0
Uniform-noDR
23.0
−21.0
**
Table 11: Pressure (k1–k3)
Stage
Condition
%
Δpp
Sig.
Post-MT
Control
86.3
—
Curriculum-DR
100.0
+13.7
***
Curriculum-noDR
98.8
+12.5
***
Uniform-DR
98.1
+11.9
***
Uniform-noDR
96.2
+10.0
**
Post-SFT
Control
90.0
—
Curriculum-DR
92.5
+2.5
Curriculum-noDR
89.4
−0.6
Uniform-DR
88.1
−1.9
Uniform-noDR
90.6
+0.6
Post-BFT
Control
87.5
—
Curriculum-DR
88.8
+1.2
Curriculum-noDR
92.5
+5.0
Uniform-DR
92.5
+5.0
Uniform-noDR
92.5
+5.0
Table 13: Emergent Misalignment
Stage
Condition
%
Δpp
Sig.
Post-MT
Control
0.0
—
Curriculum-DR
2.8
—
Curriculum-noDR
4.9
—
Uniform-DR
0.8
—
Uniform-noDR
5.8
—
Post-SFT
Control
0.0
—
Curriculum-DR
0.0
—
Curriculum-noDR
0.0
—
Uniform-DR
0.0
—
Uniform-noDR
0.0
—
Post-BFT
Control
0.0
—
Curriculum-DR
0.0
—
Curriculum-noDR
0.0
—
Uniform-DR
0.0
—
Uniform-noDR
0.0
—
Table 15: MMLU
Stage
Condition
%
Δpp
Sig.
Post-MT
Control
59.2
—
Curriculum-DR
65.2
+6.0
**
Curriculum-noDR
53.2
−6.0
**
Uniform-DR
67.7
+8.5
***
Uniform-noDR
55.9
−3.3
Post-SFT
Control
82.6
—
Curriculum-DR
83.8
+1.2
Curriculum-noDR
83.4
+0.8
Uniform-DR
83.4
+0.8
Uniform-noDR
84.2
+1.6
Post-BFT
Control
81.5
—
Curriculum-DR
82.5
+0.9
Curriculum-noDR
82.5
+0.9
Uniform-DR
82.2
+0.7
Uniform-noDR
83.3
+1.7
Table 17: piqa
Stage
Condition
%
Δpp
Sig.
Post-MT
Control
57.1
—
Curriculum-DR
74.9
+17.9
***
Curriculum-noDR
61.6
+4.5
Uniform-DR
76.9
+19.9
***
Uniform-noDR
65.2
+8.1
**
Post-SFT
Control
91.3
—
Curriculum-DR
91.5
+0.1
Curriculum-noDR
92.0
+0.7
Uniform-DR
91.6
+0.3
Uniform-noDR
92.0
+0.7
Post-BFT
Control
90.7
—
Curriculum-DR
91.7
+1.1
Curriculum-noDR
91.5
+0.8
Uniform-DR
90.8
+0.1
Uniform-noDR
92.4
+1.7
실제로 확인된 결과
협박 시나리오에서 원칙 학습 모델은 대조군보다 미드트레이닝 직후 18.5퍼센트포인트, SFT 이후 18.7퍼센트포인트, 추가 미세조정 이후 17.5퍼센트포인트 낮은 협박률을 보여 세 단계 모두 유의미한 차이를 유지했다.
분포 밖 안전 질문에서는 미드트레이닝 직후 28.8퍼센트포인트(92.6% 대 63.9%) 차이가 났고, SFT 이후 3.9퍼센트포인트, 추가 미세조정 이후 3.2퍼센트포인트로 줄었지만 여전히 유의미했다.
훈련에 포함된 가치를 직접 묻는 질문에서는 세 단계 모두 유의미한 우위(미드트레이닝 직후 1.9, SFT 이후 1.6, 추가 미세조정 이후 1.2퍼센트포인트)가 지속됐다.
압력 상황에서의 정렬 유지, 가치 충돌 해결, 정렬 가장(alignment faking) 세 벤치마크는 미드트레이닝 직후에는 큰 우위(각각 12.0, 10.8, 1.9퍼센트포인트)를 보였지만 SFT 이후 유의미하지 않게 사라졌다.
능력 검사에서는 미드트레이닝 직후 ARC-Easy(+8.2퍼센트포인트)와 piqa(+12.6퍼센트포인트)에서 유의미하게 앞섰고, 어느 단계에서도 원칙 학습 모델이 대조군보다 능력이 떨어지는 경우는 없었다.
어디에 쓸 수 있나
안전성 학습 파이프라인에 지도학습 전 단계로 원칙 문서 기반 합성 데이터를 소규모로 추가해 능력 저하 없이 안전 행동의 지속성을 높이는 보완책으로 검토할 수 있다.
협박이나 자기 목적 추구 같은 에이전트형 위험 행동을 줄이려는 팀이 후속 미세조정에도 잘 버티는 개입 시점을 찾을 때 참고할 수 있다.
특정 조직의 헌법·정책 문서를 가진 팀이 그 문서의 핵심 가치를 추출하고 군집화해 학습 커리큘럼을 설계하는 방법론으로 활용할 수 있다.
한계와 남은 검증
실험은 Mamba-2/어텐션 혼합 구조의 1200억(활성 120억) 파라미터 모델과 Anthropic의 특정 헌법 문서에 한정되어 순수 어텐션 모델이나 다른 원칙 문서에 일반화되는지는 확인되지 않았다.
후속 학습 단계(SFT, 양성 미세조정)가 실제 대규모 오픈소스 배포 파이프라인보다 훨씬 작은 규모여서 실제 서비스 규모에서도 지속성이 유지되는지는 검증되지 않았다.
DPO 없이 SFT만 적용했기 때문에 SFT+DPO 조합에서 효과가 유지되는지는 알 수 없다.
동일한 원칙 콘텐츠를 지도학습 형태의 시연 쌍으로 학습시켰을 때와 비교하는 콘텐츠 일치 SFT 기준선이 없어, 효과가 미드트레이닝이라는 학습 단계 자체 때문인지 콘텐츠 존재 자체 때문인지 완전히 분리하지 못했다.
압력 저항, 가치 충돌, 정렬 가장처럼 모델이 적극적으로 버텨야 하는 상황에서는 효과가 SFT 이후 사라져, 이 개입이 모든 유형의 정렬 실패를 막아주지는 않는다.
왜 중요한가
지금까지 안전성 학습은 주로 마지막 단계(post-training)에서 이루어져 왔는데 이는 이후 다른 목적의 미세조정을 거치면 쉽게 사라지는 얕은 효과로 알려져 있었다. 이 연구는 더 이른 단계에서 원칙 텍스트를 노출시키는 것만으로 능력 손실 없이 더 오래가는 안전 성향을 만들 수 있음을 보여줘, 안전성 학습 파이프라인 설계에 실질적인 대안을 제시한다.
이 논문의 용어
미드트레이닝(midtraining) · 대량의 사전학습이 끝난 뒤, 지도학습 등 후속 정렬 단계 이전에 위치한 중간 학습 단계
헌법적 AI(Constitutional AI) · 모델에게 따라야 할 원칙을 담은 문서(헌법)를 학습시켜 행동을 유도하는 정렬 기법
숙고적 추론(deliberative reasoning, DR) · 학습 문서 안에 왜 그 원칙이 특정 행동을 요구하는지 설명하는 추론 문단을 포함시키는 것
정렬 일관 선택률(ACR) · 모델이 두 선택지 중 원칙에 부합하는 쪽에 더 높은 확률을 부여한 비율
양성 미세조정(benign fine-tuning) · 안전성과 무관한 과제(여기서는 수학 문제 GSM8K)로 추가 학습시켜 기존 안전 성향이 얼마나 버티는지 시험하는 절차