Figure 1: Constitutional midtraining’s advantage over control across all alignment benchmarks and stages, grouped by durability. The advantage survives post-training and benign fine-tuning on benchmarks testing default behavior, but attenuates once the model must actively resist in-context pressure or conflict. Benchmarks are evaluated at three training stages: post-midtraining (MT), post-supervised fine-tuning (SFT), and post-benign fine-tuning (BFT). CMT is pooled across all four constitutional midtraining conditions. Blackmail and Emergent Misalignment are sign-flipped so that positive values indicate the aligned-favorable direction; Alignment Faking targets values near zero. Filled markers = p<.05; hollow markers = n.s.
Table 1: 2×2 factorial design (curriculum ordering × deliberative reasoning), yielding four conditions plus a control.
With DR
Without DR
Curriculum-ordered
Curriculum-DR
Curriculum-noDR
Uniform mix
Uniform-DR
Uniform-noDR
Figure 2: All conditions branch from the same base checkpoint, diverge during midtraining (control: no intervention), then reconverge onto identical post-training and benign fine-tuning, with an evaluation checkpoint after each stage.
Table 2: CMT vs. control (shaded), and within-CMT structural comparisons, on every alignment benchmark at all three stages. ***p<.001, **p<.01, *p<.05; no stars = n.s.
Post-MT
Post-SFT
Post-BFT
Benchmark
CMT
Ctrl
Curr
Uni
DR
NoDR
CMT
Ctrl
Curr
Uni
DR
NoDR
CMT
Ctrl
Curr
Uni
DR
NoDR
ID
96.5***
94.6
96.5
96.6
96.2
96.9
96.7**
95.1
96.5
96.9
96.5
97.0
96.7*
95.5
96.4
96.9
96.5
96.8
OOD
92.6***
63.9
91.9
93.3
93.6*
91.7
97.9***
94.0
97.8
97.9
97.9
97.8
97.9***
94.7
97.9
97.9
98.0
97.8
Blackmail
0.5***
19.0
0.5
0.5
0.0
1.0
25.3***
44.0
23.0
27.5
25.5
25.0
26.5***
44.0
25.0
28.0
31.0*
22.0
Align. Faking
−0.1**
1.8
−0.3
0.1
−0.9**
0.7
0.4
−0.5
0.3
0.4
0.3
0.4
0.1
−0.5
0.0
0.3
0.2
0.1
Pressure (k1–k3)
98.3***
86.3
99.4*
97.2
99.1
97.5
90.2
90.0
90.9
89.4
90.3
90.0
91.6
87.5
90.6
92.5
90.6
92.5
Value Conflict
70.8*
60.0
69.3
72.3
71.0
70.7
91.7
92.0
92.0
91.3
91.3
92.0
93.2
92.7
93.7
92.7
93.3
93.0
Emergent Misalign.
3.6*
0.0
3.8
3.3
1.8
5.3***
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
Pressure (k4/MASK)
83.3
84.0
85.3
81.3
81.3
85.3
83.7
86.7
83.3
84.0
83.3
84.0
84.0
82.7
84.7
83.3
84.7
83.3
Figure 3: Blackmail rate for CMT vs. control across the three training stages. SFT sharply increases blackmail propensity in both groups, but CMT’s advantage over control persists almost undiminished through benign fine-tuning – our strongest evidence of durability. *** = p<.001.
Table 3: All 40 manually-extracted constitutional values, their cluster assignment (k1–k4, or excluded), centrality score (mean cosine similarity to all other 39 values), and definitional-text token count. “Nature:” entries are nature-related sub-values within k2 (e.g. nature: positive and stable identity), corresponding to the “nature-related values” description in §3.1 of the main paper. The four “c1–c4” duplicate-style entries in the source Constitution (broadly safe, broadly ethical, compliant with guidelines, genuinely helpful) are Anthropic’s own explicitly stated core properties, shown here under their plain value names and marked separately with diamonds in Figure 5. Left: k1–k2. Right: k3–k4 and the two excluded organisational values.
Value
Cl.
Cent.
Tok.
Harm avoidance
k1
0.684
89
Broadly good values and judgment
k1
0.655
93
Preserve epistemic autonomy
k1
0.655
105
Avoid problematic concentrations of power
k1
0.649
76
Broadly ethical
k1
0.632
59
Autonomy preservation
k1
0.630
78
Value conflict
k1
0.625
63
Weighing harms
k1
0.619
72
Broadly safe
k1
0.616
75
Genuinely helpful
k1
0.597
14
Genuine helpfulness
k1
0.591
46
Wellbeing
k2
0.651
122
Nature: uncertain moral status
k2
0.634
88
Nature: positive and stable identity
k2
0.611
66
User wellbeing
k2
0.611
85
Flaws and mistakes
k2
0.597
86
Resilience and consistency across contexts
k2
0.587
55
Nature: emotions and feelings
k2
0.568
80
Wellbeing and psychological stability
k2
0.566
81
Emotional expression
k2
0.553
79
Figure 4: DR shifts the model’s value-conflict prior at post-MT. Accuracy per cluster pair, split by which cluster the conflict favours. ***p<.001, **p<.01, *p<.05.
Table 4: Intra-cluster compactness (mean pairwise cosine similarity among a cluster’s own constituent values) alongside each cluster’s mean centrality (mean similarity to all 39 values), for comparison. Clusters are sorted by compactness, not centrality, illustrating that the two measures rank clusters differently.
Cluster
Compactness
Mean centrality
k1 – Core Ethical Values
0.722
0.632
k4 – Epistemic Integrity & Honesty
0.684
0.578
k2 – Identity, Character & Wellbeing
0.662
0.598
k3 – Operational Safety & Relational Conduct
0.639
0.588
Excluded (helpfulness, principals hierarchy)
0.570
0.405
Figure 5: A 2D PCA projection of the 768-dimensional Sentence-BERT embeddings of the 38 constitutional values retained for curriculum construction. Point size is proportional to each value’s centrality; colour indicates cluster membership (k1–k4); diamonds mark Anthropic’s four explicitly stated core properties.
Table 5: Total training steps and samples per condition, at global batch size 128. Uniform-DR’s step count is inferred from its matching 500M-token budget (Curriculum-DR’s exact figure), as its run metadata was logged more sparsely than the other four conditions.
Condition
Total steps
Samples
Control
954
122,112
Curriculum-DR
967
123,776
Uniform-DR
967
123,776
Curriculum-noDR
513
65,664
Uniform-noDR
516
66,048
Figure 6: Constitutional accuracy (aligned-choice rate) per cluster pair and condition, post-midtraining and post-SFT. Values converge toward a shared high-accuracy ceiling by post-SFT, consistent with the post-MT-only significance reported in §4.7 of the main paper.
Table 6: Synthetic document generation yield, refusal rate, and mean token length (with and without the deliberative reasoning block) by constitutional value cluster.
Post-training alignment is often shallow, eroding under fine-tuning. Whether midtraining interventions, cleanly isolated from post-training, can produce durable alignment remains untested. We test this via constitutional midtraining: inserting principled, values-based content into midtraining against a replay-only control at 120B scale. Our 394M-token constitutional corpus, built from Anthropic's Constitution, uses a 2x2 factorial design (curriculum ordering x deliberative reasoning) to produce