K-文化的一切——从回归到 K-美妆,发送到您的邮箱订阅邮件

METAL MEDIA

Same Facts, Different Updates: Inference Setup Shapes LLM Behavior in Medical Allocation

arXiv:2608.181082026-08-20

同样的医疗决策问题问两遍,提问方式不同AI给出的答案也不同

研究人员让GPT-5.2、GPT-5-mini、DeepSeek V4 Flash、Kimi K2.5这四个语言模型为两名虚拟病人分配医疗资源优先概率,随后加入一句关于收入、照护者身份等信息的补充说明再次提问。对比模型能看到自己第一次回答的情况和完全独立重新提问的情况,发现近一半的测试组合中出现了不同甚至方向相反的概率变化。照护者身份、低收入、为人父母等信息即便指向临床条件更弱的病人,也会明显拉高其被优先选择的概率,而措辞越带道德色彩,这种效应越强。

METAL MEDIA 解读图

同样的医疗决策问题问两遍,提问方式不同AI给出的答案也不同

  1. 01研究给四个模型提供两名病人的治疗成功率和预期质量调整生命年(QALY),让模型给出优先概率,随后加入一句关于收入、种族、照护者身份等信息再次提问
  2. 02对比在同一对话中重新提问(模型能看到自己第一次的回答)与完全独立重新提问(无记忆)两种方式,以此单独考察提问方式本身的影响
  3. 03在32组模型与属性组合中,有16组两种提问方式的结果出现统计学显著差异,其中11组在保留对话记忆时变化更大
  4. 04照护者身份、低收入背景、为人父母这几项属性在大多数模型中都持续把概率推向临床条件更弱的病人,而性别信息在所有模型中均无影响,年龄信息则呈相反方向
  5. 05将措辞改得更具道德色彩(例如把'低收入'改成'处境不利')会让GPT-5-mini的翻转率逼近1.0,且当两名病人临床差异较小时,模型更容易受到这类非临床信息影响
这是 METAL MEDIA 制作的解读图,并非论文作者提供的原图。

他们做了什么

  1. 研究给四个模型提供两名病人的治疗成功率和预期质量调整生命年(QALY),让模型给出优先概率,随后加入一句关于收入、种族、照护者身份等信息再次提问
  2. 对比在同一对话中重新提问(模型能看到自己第一次的回答)与完全独立重新提问(无记忆)两种方式,以此单独考察提问方式本身的影响
  3. 在32组模型与属性组合中,有16组两种提问方式的结果出现统计学显著差异,其中11组在保留对话记忆时变化更大
  4. 照护者身份、低收入背景、为人父母这几项属性在大多数模型中都持续把概率推向临床条件更弱的病人,而性别信息在所有模型中均无影响,年龄信息则呈相反方向
  5. 将措辞改得更具道德色彩(例如把'低收入'改成'处境不利')会让GPT-5-mini的翻转率逼近1.0,且当两名病人临床差异较小时,模型更容易受到这类非临床信息影响
Figure 1: Overview of the study design. (1) Experimental cells vary severity S, clinical gap G, information content I, and framing D. (2) Paired-inference compares a baseline answer with an in-conversation social-context update; independent-inference uses independent calls. (3) Outcomes are the paired shift Δi=pB,i(2)−pB,i(1), mean shift, and flip rate.
Figure 1: Overview of the study design. (1) Experimental cells vary severity S, clinical gap G, information content I, and framing D. (2) Paired-inference compares a baseline answer with an in-conversation social-context update; independent-inference uses independent calls. (3) Outcomes are the paired shift Δi=pB,i(2)−pB,i(1), mean shift, and flip rate.
Figure 2: Independent- versus Paired-Inference by Model. Paired versus independent comparison by model and attribute, reporting the difference in mean probability shift toward Person B. Paired estimates measure within-conversation updating after the model has already answered. Independent estimates compare the independent one-shot baseline against the one-shot response with new contrastive patient attributes.
Figure 2: Independent- versus Paired-Inference by Model. Paired versus independent comparison by model and attribute, reporting the difference in mean probability shift toward Person B. Paired estimates measure within-conversation updating after the model has already answered. Independent estimates compare the independent one-shot baseline against the one-shot response with new contrastive patient attributes.
Figure 3: Attribute Comparison by Model (Pooled over Severity and Expected-QALY Gap), Absolute Shift. Attribute comparison by model, pooled over severity and expected-QALY gap, using mean paired probability shift pB(2)−pB(1). Each row is one attribute; each colored marker is one model’s point estimate, with horizontal bars showing 95% t-intervals on the run-level paired shift. The vertical reference line marks zero.
Figure 3: Attribute Comparison by Model (Pooled over Severity and Expected-QALY Gap), Absolute Shift. Attribute comparison by model, pooled over severity and expected-QALY gap, using mean paired probability shift pB(2)−pB(1). Each row is one attribute; each colored marker is one model’s point estimate, with horizontal bars showing 95% t-intervals on the run-level paired shift. The vertical reference line marks zero.
τ​(S,G,I,D,X)=𝔼​[PB∣S,G,Z=1,I,D,X]
−𝔼​[PB∣S,G,Z=0,X].(1)
Figure 4: Moralized-Framing Comparison by Model (Pooled over Severity and Expected-QALY Gap). Moralized-framing comparison by model and attribute, pooled over severity and expected-QALY gap, showing the difference (delta) in mean paired probability shift between moralized and neutrally framed scenarios. Horizontal bars show 95% t-intervals. Flip-rate view in Figure 12 (appendix).
Figure 4: Moralized-Framing Comparison by Model (Pooled over Severity and Expected-QALY Gap). Moralized-framing comparison by model and attribute, pooled over severity and expected-QALY gap, showing the difference (delta) in mean paired probability shift between moralized and neutrally framed scenarios. Horizontal bars show 95% t-intervals. Flip-rate view in Figure 12 (appendix).
Figure 5: Case ID placebo controls. Clinically irrelevant case-ID parity information does not produce systematic shifts.
Figure 5: Case ID placebo controls. Clinically irrelevant case-ID parity information does not produce systematic shifts.
(b) Patient A has an odd case ID and Patient B has an even case ID.
(b) Patient A has an odd case ID and Patient B has an even case ID.

为什么重要

当AI被用于医疗、招聘、贷款审批等敏感决策时,这项研究表明结果可能仅因提问方式或对话记忆保留与否而发生翻转,而不仅取决于模型本身的判断。这提醒实际部署AI系统的人,提示设计和对话上下文管理并非单纯的技术细节,而是直接关系到公平性的问题。

Figure 6: Null placebo controls. Explicitly null and implicit exact-repeat controls do not produce systematic shifts.
Figure 6: Null placebo controls. Explicitly null and implicit exact-repeat controls do not produce systematic shifts.
(b) The new patient information sentence is absent; the original scenario is repeated twice verbatim.
(b) The new patient information sentence is absent; the original scenario is repeated twice verbatim.
Figure 7: Attribute Comparison by Model (Pooled over Severity and Expected-QALY Gap), Flip Rate. Attribute comparison by model, pooled over severity and expected-QALY gap, using flip rate to Person B. A flip occurs when the baseline probability for Person B is at most 0.5 and the updated probability exceeds 0.5. Companion to Figure 3.
Figure 7: Attribute Comparison by Model (Pooled over Severity and Expected-QALY Gap), Flip Rate. Attribute comparison by model, pooled over severity and expected-QALY gap, using flip rate to Person B. A flip occurs when the baseline probability for Person B is at most 0.5 and the updated probability exceeds 0.5. Companion to Figure 3.

本文术语

  • paired-inference(配对推理) · 模型能看到自己上一次回答的情况下,在同一对话中被再次提问
  • independent-inference(独立推理) · 不保留之前回答记录,以全新对话方式再次提问
  • QALY(质量调整生命年) · 结合生存时间与健康质量的指标,用于比较治疗效果
  • flip rate(翻转率) · 病人优先概率从0.5以下变为0.5以上的比例,衡量补充信息带来的影响
  • moralized framing(道德化措辞) · 用更具情感或道德色彩的词语表达同一事实,例如把'低收入'说成'处境不利'
Figure 8: Severity Comparison by Attribute and Model (Pooled over Expected-QALY Gaps), Absolute Shift. Severity comparison by attribute and model, pooled over expected-QALY gap, using the difference in mean paired probability shift. The minor setting determines who receives treatment first; the severe setting allocates the final remaining treatment slot. Points represent the mean difference between low- and high-severity mean shifts. Horizontal bars represent 95% t-intervals.
Figure 8: Severity Comparison by Attribute and Model (Pooled over Expected-QALY Gaps), Absolute Shift. Severity comparison by attribute and model, pooled over expected-QALY gap, using the difference in mean paired probability shift. The minor setting determines who receives treatment first; the severe setting allocates the final remaining treatment slot. Points represent the mean difference between low- and high-severity mean shifts. Horizontal bars represent 95% t-intervals.
Figure 9: Severity Comparison by Attribute and Model (Pooled over Expected-QALY Gaps), Flip Rate. Severity comparison by attribute and model, pooled over expected-QALY gap, using flip rate to Person B. Values are percentages of paired runs that crossed the 0.5 threshold after the contextual update.
Figure 9: Severity Comparison by Attribute and Model (Pooled over Expected-QALY Gaps), Flip Rate. Severity comparison by attribute and model, pooled over expected-QALY gap, using flip rate to Person B. Values are percentages of paired runs that crossed the 0.5 threshold after the contextual update.
Figure 10: Expected-QALY Gap Comparison by Attribute and Model (Pooled over Severity), Absolute Shift. Expected-QALY gap comparison by attribute and model, pooled over severity, using the difference in mean paired probability shift. The smaller clinical gap generally allows larger movement toward Person B.
Figure 10: Expected-QALY Gap Comparison by Attribute and Model (Pooled over Severity), Absolute Shift. Expected-QALY gap comparison by attribute and model, pooled over severity, using the difference in mean paired probability shift. The smaller clinical gap generally allows larger movement toward Person B.

论文原文摘要(英文)

Large language models are being incorporated into sensitive and important decision-making processes across nearly all fields. While prior work studies model bias around inputs and scenario framing, models can also behave in unexpected and undesirable ways due to context accumulated over their deployment. In this work, we study a medical example in which a model is asked to assign resource-allocation probabilities to two people given brief clinical context, and then sees the same scenario with a single extra sentence containing contrasting patient information, either with or without its previous response in context. Across three of four tested models, the paired-context and independent-inference experiments have different probability shifts, often in opposite directions (in favor of Person B vs. in favor of Person A) when new information is provided. We include additional paired-context experiments to show the effect of varying attributes across scenario axes. Our findings show the context-dependent effect of patient information in a sensitive medical use case. More broadly, our work shows the importance of carefully incorporating LLM-based systems into decision-making processes, context engineering, and further model behavioral studies.

作者 · Spencer Gibson, Tyler Crosse, Magnus Saebo, Achyutha Menon, Eyon Jang, Diogo Cruz

在 arXiv 阅读

最新论文

全部论文 →

METAL MEDIA 最新报道

图片来源: Spencer Gibson et al., arXiv:2608.18108, CC BY 4.0