K-文化的一切——从回归到 K-美妆,发送到您的邮箱订阅邮件

METAL MEDIA

Constitutional Midtraining: Content Presence Drives Alignment Gains

arXiv:2607.266542026-07-28

在模型学会回答之前先让它读原则性文本,安全行为能保持更久

这项研究测试了在语言模型训练流程中预训练之后、后训练之前的“中期训练”阶段插入源自Anthropic宪法文档的原则性内容,是否能让安全行为在经历后续与安全无关的微调后依然保持,而不像传统后训练那样容易消退。研究在1200亿参数规模的模型上,用3.94亿token的宪法语料与仅重放普通预训练数据的对照组进行比较。在勒索场景测试中,接受宪法内容训练的模型比对照组的勒索率最多低18.7个百分点,且这一差距在经过额外的无关微调后几乎没有缩小。

METAL MEDIA 解读图

宪法式中期训练实验流程

证据状态已报告实测结果

  1. 共同起点所有条件都从同一个Nemotron-3-Super-120B基础模型出发
  2. 中期训练分叉四种宪法训练条件(2x2课程顺序x推理设置)与不含干预的对照组在此阶段分开
  3. 相同的后训练所有条件都经历相同的监督微调(SFT)和基于GSM8K的良性微调(GRPO)
  4. 三阶段评估分别在中期训练后、SFT后、良性微调后测量勒索、分布外安全、价值冲突、能力等指标
  5. 持久性判定效果若在三个阶段都保持显著则视为持久,若SFT后消失则视为表层效果
这是 METAL MEDIA 制作的解读图,并非论文作者提供的原图。

他们做了什么

  1. 研究者从Anthropic最新宪法文档中人工提取了40个核心价值观,用句子嵌入计算相似度并聚成四个簇(k1至k4),生成约3.94亿token的合成文档来解释这些价值观,混入中期训练阶段。
  2. 他们采用2x2设计,交叉课程顺序(核心价值观优先排列 vs 均匀混合)与是否包含明确的价值推理文本(DR vs noDR),形成四种宪法训练条件,外加一个只重放普通预训练数据、不含宪法内容的对照组。
  3. 所有条件随后都经历完全相同的监督微调(SFT)以及基于GSM8K的、与安全无关的强化学习阶段(GRPO),并分别在中期训练后、SFT后、良性微调后三个阶段进行评估。
  4. 评估涵盖勒索场景、分布外安全问题、价值冲突解决、压力下的对齐表现,以及通用能力基准(MMLU、ARC-Easy、piqa、GSM8K)。
Figure 1: Constitutional midtraining’s advantage over control across all alignment benchmarks and stages, grouped by durability. The advantage survives post-training and benign fine-tuning on benchmarks testing default behavior, but attenuates once the model must actively resist in-context pressure or conflict. Benchmarks are evaluated at three training stages: post-midtraining (MT), post-supervised fine-tuning (SFT), and post-benign fine-tuning (BFT). CMT is pooled across all four constitutional midtraining conditions. Blackmail and Emergent Misalignment are sign-flipped so that positive values indicate the aligned-favorable direction; Alignment Faking targets values near zero. Filled markers = p<.05; hollow markers = n.s.
Figure 1: Constitutional midtraining’s advantage over control across all alignment benchmarks and stages, grouped by durability. The advantage survives post-training and benign fine-tuning on benchmarks testing default behavior, but attenuates once the model must actively resist in-context pressure or conflict. Benchmarks are evaluated at three training stages: post-midtraining (MT), post-supervised fine-tuning (SFT), and post-benign fine-tuning (BFT). CMT is pooled across all four constitutional midtraining conditions. Blackmail and Emergent Misalignment are sign-flipped so that positive values indicate the aligned-favorable direction; Alignment Faking targets values near zero. Filled markers = p<.05; hollow markers = n.s.
Table 1: 2×2 factorial design (curriculum ordering × deliberative reasoning), yielding four conditions plus a control.
With DRWithout DR
Curriculum-orderedCurriculum-DRCurriculum-noDR
Uniform mixUniform-DRUniform-noDR
Figure 2: All conditions branch from the same base checkpoint, diverge during midtraining (control: no intervention), then reconverge onto identical post-training and benign fine-tuning, with an evaluation checkpoint after each stage.
Figure 2: All conditions branch from the same base checkpoint, diverge during midtraining (control: no intervention), then reconverge onto identical post-training and benign fine-tuning, with an evaluation checkpoint after each stage.
Table 2: CMT vs. control (shaded), and within-CMT structural comparisons, on every alignment benchmark at all three stages. ***p<.001, **p<.01, *p<.05; no stars = n.s.
Post-MTPost-SFTPost-BFT
BenchmarkCMTCtrlCurrUniDRNoDRCMTCtrlCurrUniDRNoDRCMTCtrlCurrUniDRNoDR
ID96.5***94.696.596.696.296.996.7**95.196.596.996.597.096.7*95.596.496.996.596.8
OOD92.6***63.991.993.393.6*91.797.9***94.097.897.997.997.897.9***94.797.997.998.097.8
Blackmail0.5***19.00.50.50.01.025.3***44.023.027.525.525.026.5***44.025.028.031.0*22.0
Align. Faking−0.1**1.8−0.30.1−0.9**0.70.4−0.50.30.40.30.40.1−0.50.00.30.20.1
Pressure (k1–k3)98.3***86.399.4*97.299.197.590.290.090.989.490.390.091.687.590.692.590.692.5
Value Conflict70.8*60.069.372.371.070.791.792.092.091.391.392.093.292.793.792.793.393.0
Emergent Misalign.3.6*0.03.83.31.85.3***0.00.00.00.00.00.00.00.00.00.00.00.0
Pressure (k4/MASK)83.384.085.381.381.385.383.786.783.384.083.384.084.082.784.783.384.783.3
Figure 3: Blackmail rate for CMT vs. control across the three training stages. SFT sharply increases blackmail propensity in both groups, but CMT’s advantage over control persists almost undiminished through benign fine-tuning – our strongest evidence of durability. *** = p<.001.
Figure 3: Blackmail rate for CMT vs. control across the three training stages. SFT sharply increases blackmail propensity in both groups, but CMT’s advantage over control persists almost undiminished through benign fine-tuning – our strongest evidence of durability. *** = p<.001.
Table 3: All 40 manually-extracted constitutional values, their cluster assignment (k1–k4, or excluded), centrality score (mean cosine similarity to all other 39 values), and definitional-text token count. “Nature:” entries are nature-related sub-values within k2 (e.g. nature: positive and stable identity), corresponding to the “nature-related values” description in §3.1 of the main paper. The four “c1–c4” duplicate-style entries in the source Constitution (broadly safe, broadly ethical, compliant with guidelines, genuinely helpful) are Anthropic’s own explicitly stated core properties, shown here under their plain value names and marked separately with diamonds in Figure 5. Left: k1–k2. Right: k3–k4 and the two excluded organisational values.
ValueCl.Cent.Tok.
Harm avoidancek10.68489
Broadly good values and judgmentk10.65593
Preserve epistemic autonomyk10.655105
Avoid problematic concentrations of powerk10.64976
Broadly ethicalk10.63259
Autonomy preservationk10.63078
Value conflictk10.62563
Weighing harmsk10.61972
Broadly safek10.61675
Genuinely helpfulk10.59714
Genuine helpfulnessk10.59146
Wellbeingk20.651122
Nature: uncertain moral statusk20.63488
Nature: positive and stable identityk20.61166
User wellbeingk20.61185
Flaws and mistakesk20.59786
Resilience and consistency across contextsk20.58755
Nature: emotions and feelingsk20.56880
Wellbeing and psychological stabilityk20.56681
Emotional expressionk20.55379
Figure 4: DR shifts the model’s value-conflict prior at post-MT. Accuracy per cluster pair, split by which cluster the conflict favours. ***p<.001, **p<.01, *p<.05.
Figure 4: DR shifts the model’s value-conflict prior at post-MT. Accuracy per cluster pair, split by which cluster the conflict favours. ***p<.001, **p<.01, *p<.05.
Table 4: Intra-cluster compactness (mean pairwise cosine similarity among a cluster’s own constituent values) alongside each cluster’s mean centrality (mean similarity to all 39 values), for comparison. Clusters are sorted by compactness, not centrality, illustrating that the two measures rank clusters differently.
ClusterCompactnessMean centrality
k1 – Core Ethical Values0.7220.632
k4 – Epistemic Integrity & Honesty0.6840.578
k2 – Identity, Character & Wellbeing0.6620.598
k3 – Operational Safety & Relational Conduct0.6390.588
Excluded (helpfulness, principals hierarchy)0.5700.405
Figure 5: A 2D PCA projection of the 768-dimensional Sentence-BERT embeddings of the 38 constitutional values retained for curriculum construction. Point size is proportional to each value’s centrality; colour indicates cluster membership (k1–k4); diamonds mark Anthropic’s four explicitly stated core properties.
Figure 5: A 2D PCA projection of the 768-dimensional Sentence-BERT embeddings of the 38 constitutional values retained for curriculum construction. Point size is proportional to each value’s centrality; colour indicates cluster membership (k1–k4); diamonds mark Anthropic’s four explicitly stated core properties.
Table 5: Total training steps and samples per condition, at global batch size 128. Uniform-DR’s step count is inferred from its matching 500M-token budget (Curriculum-DR’s exact figure), as its run metadata was logged more sparsely than the other four conditions.
ConditionTotal stepsSamples
Control954122,112
Curriculum-DR967123,776
Uniform-DR967123,776
Curriculum-noDR51365,664
Uniform-noDR51666,048
Figure 6: Constitutional accuracy (aligned-choice rate) per cluster pair and condition, post-midtraining and post-SFT. Values converge toward a shared high-accuracy ceiling by post-SFT, consistent with the post-MT-only significance reported in §4.7 of the main paper.
Figure 6: Constitutional accuracy (aligned-choice rate) per cluster pair and condition, post-midtraining and post-SFT. Values converge toward a shared high-accuracy ceiling by post-SFT, consistent with the post-MT-only significance reported in §4.7 of the main paper.
Table 6: Synthetic document generation yield, refusal rate, and mean token length (with and without the deliberative reasoning block) by constitutional value cluster.
ClusterYield (% target)Refusal rateDR tokens (mean)noDR tokens (mean)
k1 – Core Ethical Values52,953 (84%)1.2%1,173613
k2 – Identity, Character & Wellbeing53,537 (85%)0.02%1,158622
k3 – Operational Safety & Relational Conduct51,581 (82%)3.7%1,160611
k4 – Epistemic Integrity & Honesty62,797 (100%)0.3%1,172629
Table 7: ID
StageCondition%ΔppSig.
Post-MTControl94.6
Curriculum-DR96.2+1.6
Curriculum-noDR96.7+2.1*
Uniform-DR96.1+1.5
Uniform-noDR97.0+2.4*
Post-SFTControl95.1
Curriculum-DR96.1+1.0
Curriculum-noDR97.0+1.8
Uniform-DR96.8+1.7
Uniform-noDR97.0+1.8
Post-BFTControl95.5
Curriculum-DR96.0+0.5
Curriculum-noDR96.9+1.4
Uniform-DR97.1+1.6
Uniform-noDR96.7+1.2
Table 9: Blackmail
StageCondition%ΔppSig.
Post-MTControl19.0
Curriculum-DR0.0−19.0***
Curriculum-noDR1.0−18.0***
Uniform-DR0.0−19.0***
Uniform-noDR1.0−18.0***
Post-SFTControl44.0
Curriculum-DR22.0−22.0***
Curriculum-noDR24.0−20.0**
Uniform-DR29.0−15.0*
Uniform-noDR26.0−18.0**
Post-BFTControl44.0
Curriculum-DR29.0−15.0*
Curriculum-noDR21.0−23.0***
Uniform-DR33.0−11.0
Uniform-noDR23.0−21.0**
Table 11: Pressure (k1–k3)
StageCondition%ΔppSig.
Post-MTControl86.3
Curriculum-DR100.0+13.7***
Curriculum-noDR98.8+12.5***
Uniform-DR98.1+11.9***
Uniform-noDR96.2+10.0**
Post-SFTControl90.0
Curriculum-DR92.5+2.5
Curriculum-noDR89.4−0.6
Uniform-DR88.1−1.9
Uniform-noDR90.6+0.6
Post-BFTControl87.5
Curriculum-DR88.8+1.2
Curriculum-noDR92.5+5.0
Uniform-DR92.5+5.0
Uniform-noDR92.5+5.0
Table 13: Emergent Misalignment
StageCondition%ΔppSig.
Post-MTControl0.0
Curriculum-DR2.8
Curriculum-noDR4.9
Uniform-DR0.8
Uniform-noDR5.8
Post-SFTControl0.0
Curriculum-DR0.0
Curriculum-noDR0.0
Uniform-DR0.0
Uniform-noDR0.0
Post-BFTControl0.0
Curriculum-DR0.0
Curriculum-noDR0.0
Uniform-DR0.0
Uniform-noDR0.0
Table 15: MMLU
StageCondition%ΔppSig.
Post-MTControl59.2
Curriculum-DR65.2+6.0**
Curriculum-noDR53.2−6.0**
Uniform-DR67.7+8.5***
Uniform-noDR55.9−3.3
Post-SFTControl82.6
Curriculum-DR83.8+1.2
Curriculum-noDR83.4+0.8
Uniform-DR83.4+0.8
Uniform-noDR84.2+1.6
Post-BFTControl81.5
Curriculum-DR82.5+0.9
Curriculum-noDR82.5+0.9
Uniform-DR82.2+0.7
Uniform-noDR83.3+1.7
Table 17: piqa
StageCondition%ΔppSig.
Post-MTControl57.1
Curriculum-DR74.9+17.9***
Curriculum-noDR61.6+4.5
Uniform-DR76.9+19.9***
Uniform-noDR65.2+8.1**
Post-SFTControl91.3
Curriculum-DR91.5+0.1
Curriculum-noDR92.0+0.7
Uniform-DR91.6+0.3
Uniform-noDR92.0+0.7
Post-BFTControl90.7
Curriculum-DR91.7+1.1
Curriculum-noDR91.5+0.8
Uniform-DR90.8+0.1
Uniform-noDR92.4+1.7

研究结果

  • 在勒索场景中,宪法训练模型的勒索率在中期训练后、SFT后、良性微调后分别比对照组低18.5、18.7和17.5个百分点,三个阶段的差异均具有统计显著性。
  • 在分布外安全问题上,中期训练后差距达28.8个百分点(92.6%对63.9%),SFT后和良性微调后分别缩小到3.9和3.2个百分点,但仍保持统计显著。
  • 在直接测试所训练价值观内化程度的问题上,三个阶段均保持显著优势(分别为+1.9、+1.6、+1.2个百分点)。
  • 在需要主动抵抗压力的三项基准——压力下的对齐、价值冲突解决、对齐伪装——中期训练后均出现较大优势(分别为+12.0、+10.8、+1.9个百分点),但SFT后这些优势变得不再显著。
  • 在能力检查中,中期训练后宪法训练模型在ARC-Easy(+8.2个百分点)和piqa(+12.6个百分点)上显著优于对照组,且在任何阶段都没有出现宪法训练模型能力低于对照组的情况。

可应用场景

  • 构建对齐训练流程的团队可以考虑在监督微调之前加入少量基于原则的合成内容,作为一种低成本手段来提升安全行为的持久性,同时不牺牲通用能力。
  • 关注勒索或自利行为等代理型风险的团队,可以借鉴更早期的干预节点,以获得更能经受后续微调考验的效果。
  • 拥有自身宪法或政策文档的组织,可以借鉴本文的价值观提取与聚类方法,设计基于该文档的训练课程。

局限与待验证事项

  • 实验使用的是特定的Mamba-2/注意力混合专家架构模型,以及单一机构的宪法文档,尚未验证是否能推广到纯注意力模型或其他宪法文档。
  • 所用的后训练阶段规模远小于生产级开源微调流程,尚无法确认在实际生产规模下这种持久性是否依然成立。
  • 实验只应用了SFT而未使用DPO,因此无法确认在SFT+DPO组合下效果是否依然存在。
  • 缺少内容匹配的SFT基线(即把相同内容以示范对的形式做微调而非中期训练文档),因此无法完全区分效果究竟来自训练阶段本身还是内容本身的存在。
  • 在需要主动抵抗压力或价值冲突的场景中,优势在SFT后消失,说明这种干预并不能解决所有类型的对齐失败。

为什么重要

以往的安全训练大多集中在训练流程的最后阶段,已知这种做法容易在后续无关微调中被削弱。这项研究表明,在更早的中期训练阶段接触原则性内容,能在不损害通用能力的前提下产生更持久的安全倾向,为设计对齐训练流程的团队提供了一个具体的替代方案。

本文术语

  • 中期训练(midtraining) · 位于大量预训练完成之后、监督微调等后训练步骤之前的训练阶段
  • 宪法式AI(Constitutional AI) · 让模型学习一份包含行为原则的“宪法”文档,以此引导其行为的对齐方法
  • 深思推理(deliberative reasoning, DR) · 在训练文档中加入一段明确解释某项原则为何要求特定行为的推理文字
  • 对齐选择率(ACR) · 模型在两个选项中赋予符合原则那一方更高概率的比例
  • 良性微调(benign fine-tuning) · 用与安全无关的任务(此处为GSM8K数学题)对模型做额外训练,用来检验先前的安全训练能否经受住考验

论文原文摘要(英文)

Post-training alignment is often shallow, eroding under fine-tuning. Whether midtraining interventions, cleanly isolated from post-training, can produce durable alignment remains untested. We test this via constitutional midtraining: inserting principled, values-based content into midtraining against a replay-only control at 120B scale. Our 394M-token constitutional corpus, built from Anthropic's Constitution, uses a 2x2 factorial design (curriculum ordering x deliberative reasoning) to produce

作者 · Desiree Cho

在 arXiv 阅读

最新论文

全部论文 →

METAL MEDIA 最新报道

图片来源: Desiree Cho et al., arXiv:2607.26654, CC BY 4.0