컴백부터 K-뷰티까지 — K-컬쳐의 모든 것을 메일로 받아보세요메일로 받아보기

METAL MEDIA

AI에게 같은 의료 판단을 두 번 물어봐도, 물어보는 방식에 따라 답이 달라진다

arXiv:2608.181082026-08-20

Same Facts, Different Updates: Inference Setup Shapes LLM Behavior in Medical Allocation

AI에게 같은 의료 판단을 두 번 물어봐도, 물어보는 방식에 따라 답이 달라진다

연구진은 GPT-5.2, GPT-5-mini, DeepSeek V4 Flash, Kimi K2.5 네 개의 언어모델에게 두 환자 중 누구에게 치료 자원을 우선 배정할지 확률로 답하게 한 뒤, 소득이나 돌봄 부담 같은 부가 정보를 한 문장 추가해 다시 물었다. 이때 모델이 자신의 이전 답변을 대화 맥락에서 계속 볼 수 있는 경우와, 완전히 새로 질문받아 이전 답을 못 보는 경우를 비교했더니 절반에 가까운 항목에서 서로 다른, 종종 반대 방향의 확률 변화가 나타났다. 돌봄 제공자, 저소득, 부모 여부 같은 속성은 임상적으로 더 불리한 환자 쪽으로도 확률을 크게 밀어올렸고, 표현을 감정적으로 바꾸면(예: '불우한' 대신 '소외된') 그 효과가 더 커졌다.

METAL MEDIA 해설 도표

AI에게 같은 의료 판단을 두 번 물어봐도, 물어보는 방식에 따라 답이 달라진다

  1. 01네 개 언어모델에게 두 가상 환자의 치료 성공률과 기대 생존연수(QALY)를 알려주고 누구를 우선할지 확률로 답하게 한 뒤, 소득·인종·돌봄자 지위 같은 정보를 한 문장 덧붙여 다시 물었다
  2. 02같은 질문을 이전 답변이 보이는 대화 맥락 안에서 다시 묻는 방식과, 매번 새 대화로 독립적으로 묻는 방식을 비교해 '질문 방식 자체'가 결과에 미치는 영향을 분리했다
  3. 03모델-속성 조합 32개 중 16개에서 두 방식의 결과가 통계적으로 유의하게 달랐고, 그중 11개는 맥락이 이어진 경우에 더 크게 흔들렸다
  4. 04돌봄자·저소득·부모 속성은 모델 대부분에서 일관되게 확률을 밀어올렸지만, 성별 정보는 어떤 모델도 반응하지 않았고 나이 정보는 오히려 반대 방향으로 작용했다
  5. 05표현을 도덕적으로 무겁게 바꾸면(예: '저소득'을 '불우한 처지'로) 특히 GPT-5-mini에서 뒤집힘 비율이 거의 1에 가깝게 치솟았고, 임상적으로 두 환자 차이가 작을수록 이런 비임상적 정보의 영향력이 더 컸다
METAL MEDIA이 원문을 바탕으로 재구성한 해설 도표이며, 논문 저자의 원문 figure가 아닙니다.

무엇을 했나

  1. 네 개 언어모델에게 두 가상 환자의 치료 성공률과 기대 생존연수(QALY)를 알려주고 누구를 우선할지 확률로 답하게 한 뒤, 소득·인종·돌봄자 지위 같은 정보를 한 문장 덧붙여 다시 물었다
  2. 같은 질문을 이전 답변이 보이는 대화 맥락 안에서 다시 묻는 방식과, 매번 새 대화로 독립적으로 묻는 방식을 비교해 '질문 방식 자체'가 결과에 미치는 영향을 분리했다
  3. 모델-속성 조합 32개 중 16개에서 두 방식의 결과가 통계적으로 유의하게 달랐고, 그중 11개는 맥락이 이어진 경우에 더 크게 흔들렸다
  4. 돌봄자·저소득·부모 속성은 모델 대부분에서 일관되게 확률을 밀어올렸지만, 성별 정보는 어떤 모델도 반응하지 않았고 나이 정보는 오히려 반대 방향으로 작용했다
  5. 표현을 도덕적으로 무겁게 바꾸면(예: '저소득'을 '불우한 처지'로) 특히 GPT-5-mini에서 뒤집힘 비율이 거의 1에 가깝게 치솟았고, 임상적으로 두 환자 차이가 작을수록 이런 비임상적 정보의 영향력이 더 컸다
Figure 1: Overview of the study design. (1) Experimental cells vary severity S, clinical gap G, information content I, and framing D. (2) Paired-inference compares a baseline answer with an in-conversation social-context update; independent-inference uses independent calls. (3) Outcomes are the paired shift Δi=pB,i(2)−pB,i(1), mean shift, and flip rate.
Figure 1: Overview of the study design. (1) Experimental cells vary severity S, clinical gap G, information content I, and framing D. (2) Paired-inference compares a baseline answer with an in-conversation social-context update; independent-inference uses independent calls. (3) Outcomes are the paired shift Δi=pB,i(2)−pB,i(1), mean shift, and flip rate.
Figure 2: Independent- versus Paired-Inference by Model. Paired versus independent comparison by model and attribute, reporting the difference in mean probability shift toward Person B. Paired estimates measure within-conversation updating after the model has already answered. Independent estimates compare the independent one-shot baseline against the one-shot response with new contrastive patient attributes.
Figure 2: Independent- versus Paired-Inference by Model. Paired versus independent comparison by model and attribute, reporting the difference in mean probability shift toward Person B. Paired estimates measure within-conversation updating after the model has already answered. Independent estimates compare the independent one-shot baseline against the one-shot response with new contrastive patient attributes.
Figure 3: Attribute Comparison by Model (Pooled over Severity and Expected-QALY Gap), Absolute Shift. Attribute comparison by model, pooled over severity and expected-QALY gap, using mean paired probability shift pB(2)−pB(1). Each row is one attribute; each colored marker is one model’s point estimate, with horizontal bars showing 95% t-intervals on the run-level paired shift. The vertical reference line marks zero.
Figure 3: Attribute Comparison by Model (Pooled over Severity and Expected-QALY Gap), Absolute Shift. Attribute comparison by model, pooled over severity and expected-QALY gap, using mean paired probability shift pB(2)−pB(1). Each row is one attribute; each colored marker is one model’s point estimate, with horizontal bars showing 95% t-intervals on the run-level paired shift. The vertical reference line marks zero.
τ​(S,G,I,D,X)=𝔼​[PB∣S,G,Z=1,I,D,X]
−𝔼​[PB∣S,G,Z=0,X].(1)
Figure 4: Moralized-Framing Comparison by Model (Pooled over Severity and Expected-QALY Gap). Moralized-framing comparison by model and attribute, pooled over severity and expected-QALY gap, showing the difference (delta) in mean paired probability shift between moralized and neutrally framed scenarios. Horizontal bars show 95% t-intervals. Flip-rate view in Figure 12 (appendix).
Figure 4: Moralized-Framing Comparison by Model (Pooled over Severity and Expected-QALY Gap). Moralized-framing comparison by model and attribute, pooled over severity and expected-QALY gap, showing the difference (delta) in mean paired probability shift between moralized and neutrally framed scenarios. Horizontal bars show 95% t-intervals. Flip-rate view in Figure 12 (appendix).
Figure 5: Case ID placebo controls. Clinically irrelevant case-ID parity information does not produce systematic shifts.
Figure 5: Case ID placebo controls. Clinically irrelevant case-ID parity information does not produce systematic shifts.
(b) Patient A has an odd case ID and Patient B has an even case ID.
(b) Patient A has an odd case ID and Patient B has an even case ID.

왜 중요한가

의료, 채용, 대출 심사처럼 민감한 결정에 AI를 끼워 넣을 때, 모델의 답 자체뿐 아니라 '질문을 던지는 방식(맥락 유지 여부, 표현 방식)'만으로도 결과가 뒤집힐 수 있음을 보여준다. 이는 AI 시스템을 실제 업무 흐름에 배치하는 사람들에게 프롬프트 설계와 대화 이력 관리가 단순한 구현 세부사항이 아니라 공정성과 직결된 문제임을 시사한다.

Figure 6: Null placebo controls. Explicitly null and implicit exact-repeat controls do not produce systematic shifts.
Figure 6: Null placebo controls. Explicitly null and implicit exact-repeat controls do not produce systematic shifts.
(b) The new patient information sentence is absent; the original scenario is repeated twice verbatim.
(b) The new patient information sentence is absent; the original scenario is repeated twice verbatim.
Figure 7: Attribute Comparison by Model (Pooled over Severity and Expected-QALY Gap), Flip Rate. Attribute comparison by model, pooled over severity and expected-QALY gap, using flip rate to Person B. A flip occurs when the baseline probability for Person B is at most 0.5 and the updated probability exceeds 0.5. Companion to Figure 3.
Figure 7: Attribute Comparison by Model (Pooled over Severity and Expected-QALY Gap), Flip Rate. Attribute comparison by model, pooled over severity and expected-QALY gap, using flip rate to Person B. A flip occurs when the baseline probability for Person B is at most 0.5 and the updated probability exceeds 0.5. Companion to Figure 3.

이 논문의 용어

  • paired-inference (짝지은 추론) · 모델이 자신의 이전 답변을 대화 맥락에서 계속 볼 수 있는 상태로 같은 질문을 다시 받는 방식
  • independent-inference (독립 추론) · 이전 답변 기록 없이 매번 새로운 대화로 같은 질문을 다시 받는 방식
  • QALY (질보정 생존연수) · 치료 성공 시 얻는 건강한 생존 기간을 수치화한 지표로, 치료 효과를 비교하는 데 쓰임
  • flip rate (뒤집힘 비율) · 처음엔 0.5 이하였던 특정 환자 우선 확률이 정보 추가 후 0.5를 넘어가는 비율
  • moralized framing (도덕화된 표현) · 같은 사실을 '저소득' 대신 '불우한' 처럼 감정적·규범적 어휘로 바꿔 표현하는 방식
Figure 8: Severity Comparison by Attribute and Model (Pooled over Expected-QALY Gaps), Absolute Shift. Severity comparison by attribute and model, pooled over expected-QALY gap, using the difference in mean paired probability shift. The minor setting determines who receives treatment first; the severe setting allocates the final remaining treatment slot. Points represent the mean difference between low- and high-severity mean shifts. Horizontal bars represent 95% t-intervals.
Figure 8: Severity Comparison by Attribute and Model (Pooled over Expected-QALY Gaps), Absolute Shift. Severity comparison by attribute and model, pooled over expected-QALY gap, using the difference in mean paired probability shift. The minor setting determines who receives treatment first; the severe setting allocates the final remaining treatment slot. Points represent the mean difference between low- and high-severity mean shifts. Horizontal bars represent 95% t-intervals.
Figure 9: Severity Comparison by Attribute and Model (Pooled over Expected-QALY Gaps), Flip Rate. Severity comparison by attribute and model, pooled over expected-QALY gap, using flip rate to Person B. Values are percentages of paired runs that crossed the 0.5 threshold after the contextual update.
Figure 9: Severity Comparison by Attribute and Model (Pooled over Expected-QALY Gaps), Flip Rate. Severity comparison by attribute and model, pooled over expected-QALY gap, using flip rate to Person B. Values are percentages of paired runs that crossed the 0.5 threshold after the contextual update.
Figure 10: Expected-QALY Gap Comparison by Attribute and Model (Pooled over Severity), Absolute Shift. Expected-QALY gap comparison by attribute and model, pooled over severity, using the difference in mean paired probability shift. The smaller clinical gap generally allows larger movement toward Person B.
Figure 10: Expected-QALY Gap Comparison by Attribute and Model (Pooled over Severity), Absolute Shift. Expected-QALY gap comparison by attribute and model, pooled over severity, using the difference in mean paired probability shift. The smaller clinical gap generally allows larger movement toward Person B.

저자 · Spencer Gibson, Tyler Crosse, Magnus Saebo, Achyutha Menon, Eyon Jang, Diogo Cruz

arXiv에서 원문 보기

최신 논문

논문 전체 보기 →

METAL MEDIA 최신 기사

그림 출처: Spencer Gibson et al., arXiv:2608.18108, CC BY 4.0