Same Facts, Different Updates: Inference Setup Shapes LLM Behavior in Medical Allocation
AI에게 같은 의료 판단을 두 번 물어봐도, 물어보는 방식에 따라 답이 달라진다
연구진은 GPT-5.2, GPT-5-mini, DeepSeek V4 Flash, Kimi K2.5 네 개의 언어모델에게 두 환자 중 누구에게 치료 자원을 우선 배정할지 확률로 답하게 한 뒤, 소득이나 돌봄 부담 같은 부가 정보를 한 문장 추가해 다시 물었다. 이때 모델이 자신의 이전 답변을 대화 맥락에서 계속 볼 수 있는 경우와, 완전히 새로 질문받아 이전 답을 못 보는 경우를 비교했더니 절반에 가까운 항목에서 서로 다른, 종종 반대 방향의 확률 변화가 나타났다. 돌봄 제공자, 저소득, 부모 여부 같은 속성은 임상적으로 더 불리한 환자 쪽으로도 확률을 크게 밀어올렸고, 표현을 감정적으로 바꾸면(예: '불우한' 대신 '소외된') 그 효과가 더 커졌다.
METAL MEDIA 해설 도표
AI에게 같은 의료 판단을 두 번 물어봐도, 물어보는 방식에 따라 답이 달라진다
01네 개 언어모델에게 두 가상 환자의 치료 성공률과 기대 생존연수(QALY)를 알려주고 누구를 우선할지 확률로 답하게 한 뒤, 소득·인종·돌봄자 지위 같은 정보를 한 문장 덧붙여 다시 물었다
02같은 질문을 이전 답변이 보이는 대화 맥락 안에서 다시 묻는 방식과, 매번 새 대화로 독립적으로 묻는 방식을 비교해 '질문 방식 자체'가 결과에 미치는 영향을 분리했다
03모델-속성 조합 32개 중 16개에서 두 방식의 결과가 통계적으로 유의하게 달랐고, 그중 11개는 맥락이 이어진 경우에 더 크게 흔들렸다
04돌봄자·저소득·부모 속성은 모델 대부분에서 일관되게 확률을 밀어올렸지만, 성별 정보는 어떤 모델도 반응하지 않았고 나이 정보는 오히려 반대 방향으로 작용했다
05표현을 도덕적으로 무겁게 바꾸면(예: '저소득'을 '불우한 처지'로) 특히 GPT-5-mini에서 뒤집힘 비율이 거의 1에 가깝게 치솟았고, 임상적으로 두 환자 차이가 작을수록 이런 비임상적 정보의 영향력이 더 컸다
METAL MEDIA이 원문을 바탕으로 재구성한 해설 도표이며, 논문 저자의 원문 figure가 아닙니다.
무엇을 했나
네 개 언어모델에게 두 가상 환자의 치료 성공률과 기대 생존연수(QALY)를 알려주고 누구를 우선할지 확률로 답하게 한 뒤, 소득·인종·돌봄자 지위 같은 정보를 한 문장 덧붙여 다시 물었다
같은 질문을 이전 답변이 보이는 대화 맥락 안에서 다시 묻는 방식과, 매번 새 대화로 독립적으로 묻는 방식을 비교해 '질문 방식 자체'가 결과에 미치는 영향을 분리했다
모델-속성 조합 32개 중 16개에서 두 방식의 결과가 통계적으로 유의하게 달랐고, 그중 11개는 맥락이 이어진 경우에 더 크게 흔들렸다
돌봄자·저소득·부모 속성은 모델 대부분에서 일관되게 확률을 밀어올렸지만, 성별 정보는 어떤 모델도 반응하지 않았고 나이 정보는 오히려 반대 방향으로 작용했다
표현을 도덕적으로 무겁게 바꾸면(예: '저소득'을 '불우한 처지'로) 특히 GPT-5-mini에서 뒤집힘 비율이 거의 1에 가깝게 치솟았고, 임상적으로 두 환자 차이가 작을수록 이런 비임상적 정보의 영향력이 더 컸다
Figure 1: Overview of the study design. (1) Experimental cells vary severity S, clinical gap G, information content I, and framing D. (2) Paired-inference compares a baseline answer with an in-conversation social-context update; independent-inference uses independent calls. (3) Outcomes are the paired shift Δi=pB,i(2)−pB,i(1), mean shift, and flip rate.Figure 2: Independent- versus Paired-Inference by Model. Paired versus independent comparison by model and attribute, reporting the difference in mean probability shift toward Person B. Paired estimates measure within-conversation updating after the model has already answered. Independent estimates compare the independent one-shot baseline against the one-shot response with new contrastive patient attributes.Figure 3: Attribute Comparison by Model (Pooled over Severity and Expected-QALY Gap), Absolute Shift. Attribute comparison by model, pooled over severity and expected-QALY gap, using mean paired probability shift pB(2)−pB(1). Each row is one attribute; each colored marker is one model’s point estimate, with horizontal bars showing 95% t-intervals on the run-level paired shift. The vertical reference line marks zero.
τ(S,G,I,D,X)
=𝔼[PB∣S,G,Z=1,I,D,X]
−𝔼[PB∣S,G,Z=0,X].
(1)
Figure 4: Moralized-Framing Comparison by Model (Pooled over Severity and Expected-QALY Gap). Moralized-framing comparison by model and attribute, pooled over severity and expected-QALY gap, showing the difference (delta) in mean paired probability shift between moralized and neutrally framed scenarios. Horizontal bars show 95% t-intervals. Flip-rate view in Figure 12 (appendix).Figure 5: Case ID placebo controls. Clinically irrelevant case-ID parity information does not produce systematic shifts.(b) Patient A has an odd case ID and Patient B has an even case ID.
왜 중요한가
의료, 채용, 대출 심사처럼 민감한 결정에 AI를 끼워 넣을 때, 모델의 답 자체뿐 아니라 '질문을 던지는 방식(맥락 유지 여부, 표현 방식)'만으로도 결과가 뒤집힐 수 있음을 보여준다. 이는 AI 시스템을 실제 업무 흐름에 배치하는 사람들에게 프롬프트 설계와 대화 이력 관리가 단순한 구현 세부사항이 아니라 공정성과 직결된 문제임을 시사한다.
Figure 6: Null placebo controls. Explicitly null and implicit exact-repeat controls do not produce systematic shifts.(b) The new patient information sentence is absent; the original scenario is repeated twice verbatim.Figure 7: Attribute Comparison by Model (Pooled over Severity and Expected-QALY Gap), Flip Rate. Attribute comparison by model, pooled over severity and expected-QALY gap, using flip rate to Person B. A flip occurs when the baseline probability for Person B is at most 0.5 and the updated probability exceeds 0.5. Companion to Figure 3.
이 논문의 용어
paired-inference (짝지은 추론) · 모델이 자신의 이전 답변을 대화 맥락에서 계속 볼 수 있는 상태로 같은 질문을 다시 받는 방식
independent-inference (독립 추론) · 이전 답변 기록 없이 매번 새로운 대화로 같은 질문을 다시 받는 방식
QALY (질보정 생존연수) · 치료 성공 시 얻는 건강한 생존 기간을 수치화한 지표로, 치료 효과를 비교하는 데 쓰임
flip rate (뒤집힘 비율) · 처음엔 0.5 이하였던 특정 환자 우선 확률이 정보 추가 후 0.5를 넘어가는 비율
moralized framing (도덕화된 표현) · 같은 사실을 '저소득' 대신 '불우한' 처럼 감정적·규범적 어휘로 바꿔 표현하는 방식
Figure 8: Severity Comparison by Attribute and Model (Pooled over Expected-QALY Gaps), Absolute Shift. Severity comparison by attribute and model, pooled over expected-QALY gap, using the difference in mean paired probability shift. The minor setting determines who receives treatment first; the severe setting allocates the final remaining treatment slot. Points represent the mean difference between low- and high-severity mean shifts. Horizontal bars represent 95% t-intervals.Figure 9: Severity Comparison by Attribute and Model (Pooled over Expected-QALY Gaps), Flip Rate. Severity comparison by attribute and model, pooled over expected-QALY gap, using flip rate to Person B. Values are percentages of paired runs that crossed the 0.5 threshold after the contextual update.Figure 10: Expected-QALY Gap Comparison by Attribute and Model (Pooled over Severity), Absolute Shift. Expected-QALY gap comparison by attribute and model, pooled over severity, using the difference in mean paired probability shift. The smaller clinical gap generally allows larger movement toward Person B.