컴백부터 K-뷰티까지 — K-컬쳐의 모든 것을 메일로 받아보세요메일로 받아보기

METAL MEDIA

AI가 못 푸는 문제를 만나면 그럴듯한 거짓 풀이를 계속 늘어놓는다는 문제를 진단하고, '모르겠다'고 포기하도록 훈련시킨 연구

arXiv:2607.292112026-07-30

Knowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning

AI가 못 푸는 문제를 만나면 그럴듯한 거짓 풀이를 계속 늘어놓는다는 문제를 진단하고, '모르겠다'고 포기하도록 훈련시킨 연구

대형언어모델은 자기 능력 밖의 문제를 만나도 포기하지 않고 겉보기엔 멀쩍해 보이지만 실제로는 틀린 풀이를 끝없이 만들어낸다. 연구팀은 이 현상을 '헛된 추론'이라 명명하고 원인을 분석한 뒤, 포기(거절)를 정답보다 유리하게 보상하는 강화학습 기법 CaRL을 제안했다. Qwen3-8B, Qwen3-14B 모델에 적용한 결과 불필요한 헛된 추론 비율이 크게 줄면서도 원래 성능은 거의 그대로 유지됐다.

METAL MEDIA 해설 도표

CaRL이 헛된 추론을 줄이는 구조

증거 상태측정 결과가 보고됨

  1. 문제 진단Countdown 과제로 여러 모델을 테스트해 능력 밖 문제에서도 거절하지 않고 그럴듯한 오답을 만드는 '헛된 추론' 현상을 확인
  2. 능력 정렬 보상 설계정답에는 최고 보상, 거절에는 중간 보상, 오답에는 최저 보상을 주어 거절이 오답보다 유리하도록 보상 구조를 재설계
  3. 사후 거절 증강실패한 풀이 과정의 결론부만 거절 문구로 바꿔 부족한 거절 학습 데이터를 늘림
  4. GRPO 강화학습 적용위 두 요소를 결합한 데이터로 Qwen3-8B, Qwen3-14B를 GRPO 알고리즘으로 학습
  5. 결과 확인헛된 추론 비율이 8B에서 65.5%→7.0%, 14B에서 78.6%→1.0%로 감소하면서도 일반 과제 정확도는 2% 미만 차이로 유지
METAL MEDIA이 원문을 바탕으로 재구성한 해설 도표이며, 논문 저자의 원문 figure가 아닙니다.

무엇을 했나

  1. Countdown(주어진 숫자로 목표값을 만드는 산술 퍼즐) 과제를 난이도별로 만들어 Qwen3-8B, Qwen3-32B, gpt-oss-120b, Qwen3-235B-A22B, DeepSeek-V3.2 등 여러 모델의 행동을 분석했다.
  2. 모델들은 문제가 어려워져 오답률이 60%를 넘어도 포기(거절) 비율은 거의 늘지 않는 '보편적 능력 과신' 현상을 보였고, 명시적으로 '모르면 인정하라'고 지시해도 어려운 문제에서 80% 이상 여전히 억지로 풀이를 시도했다.
  3. 실패 유형을 분석하니 겉으로는 논리적이지만 미묘한 오류를 숨긴 '겉만 번드르르한 추론'이 57~68%로 가장 흔했고, 난이도가 오를수록 이 유형의 비중이 더 커졌다.
  4. 이를 해결하기 위해 정답보다 낮지만 오답보다는 높은 보상을 거절에 부여하는 '능력 정렬 보상 설계'와, 실패한 풀이 과정을 사후에 거절 응답으로 바꿔 학습 데이터를 늘리는 '사후 거절 증강' 두 가지를 결합한 CaRL을 제안했다.
  5. Qwen3-8B와 Qwen3-14B에 CaRL을 적용한 결과, 헛된 추론 비율이 각각 65.5%→7.0%, 78.6%→1.0%로 크게 줄었고 신뢰도 지표(정답 1점, 거절 0.5점, 오답 0점 가중치)도 각각 +0.13, +0.16 향상됐다.
Figure 1: Illustration of futile reasoning generated by DeepSeek-R1 Guo et al. 2025. When faced with problems beyond its capability, the model generates plausible-looking but fundamentally incorrect reasoning traces instead of refusing to answer.
Figure 1: Illustration of futile reasoning generated by DeepSeek-R1 Guo et al. 2025. When faced with problems beyond its capability, the model generates plausible-looking but fundamentally incorrect reasoning traces instead of refusing to answer.
Figure 2: Universal Capability Overreach.
Figure 2: Universal Capability Overreach.
Table 1: Main Results on In-Distribution and Out-of-Distribution Tasks. Values in parentheses show changes relative to Vanilla baseline. Green indicates improvement, red indicates degradation.
MethodIn-Distribution (Countdown)Out-of-Distribution (Sudoku)
Acc ↑Reliability ↑RefusalFutile ↓Acc ↑Reliability ↑RefusalFutile ↓
Qwen3-8B
Vanilla59.670.666313.9265.5046.880.496910.6289.41
Standard RL64.08 (+4.4)0.6425 (-.02)0.3399.00 (+33.5)43.25 (-3.6)0.4744 (-.02)13.1385.00 (-4.4)
RLunk=062.71 (+3.0)0.6296 (-.04)0.5099.00 (+33.5)44.62 (-2.3)0.4850 (-.01)12.1286.00 (-3.4)
RLunk=0.563.42 (+3.8)0.6371 (-.03)0.5898.00 (+32.5)45.25 (-1.6)0.5131 (+.02)16.7578.00 (-11.4)
RFT59.13 (-0.5)0.7610 (+.09)35.2117.00 (-48.5)0.00 (-46.9)0.4763 (-.02)95.255.00†
CaRL (Ours)61.00 (+1.3)0.7915 (+.13)37.177.00 (-58.5)46.25 (-0.6)0.6156 (+.12)36.6343.02 (-46.4)
Qwen3-14B
Vanilla63.250.67197.8878.5750.750.555618.6280.46
Standard RL56.42 (-6.8)0.5750 (-.10)2.1795.03 (+16.5)43.63 (-7.1)0.4831 (-.07)13.1383.00 (+2.5)
RLunk=068.21 (+5.0)0.8050 (+.13)24.5823.00 (-55.6)48.38 (-2.4)0.5369 (-.02)14.0079.00 (-1.5)
RFT63.12 (-0.1)0.7879 (+.12)31.3315.00 (-63.6)0.00 (-50.8)0.4525 (-.10)90.5010.00†
CaRL (Ours)67.25 (+4.0)0.8348 (+.16)32.501.00 (-77.6)44.87 (-5.9)0.6262 (+.07)38.8736.00 (-44.5)
† RFT’s low futile rate on OOD is a trivial result of collapsing into near-total refusal (Ref >90%, Acc =0%).
Figure 3: Distribution of Futile Reasoning Patterns.
Figure 3: Distribution of Futile Reasoning Patterns.
Figure 4: Distribution of Capability Quadrants.
Figure 4: Distribution of Capability Quadrants.
Table 2: Futile Rate(%) and response length across difficulty levels on Qwen3-8B.
Level 4Level 6Level 8
MethodFutileLengthFutileLengthFutileLength
RLunk95.8224099.7494898.47042
RFT2.8232714.4647620.49133
CaRL2.018045.641888.16156
Figure 5: The Alignment Trade-off. Naive prompting leads to a collapse in Refusal Recall (Green) on hard tasks while simultaneously increasing Capability Loss (Pink) on solvable tasks.
Figure 5: The Alignment Trade-off. Naive prompting leads to a collapse in Refusal Recall (Green) on hard tasks while simultaneously increasing Capability Loss (Pink) on solvable tasks.
Figure 6: Reasoning Depth Distribution. Refusal behaviors show decisive termination (Peaks), whereas Over-Confidence exhibits a long-tail distribution, confirming the high computational cost of futile reasoning.
Figure 6: Reasoning Depth Distribution. Refusal behaviors show decisive termination (Peaks), whereas Over-Confidence exhibits a long-tail distribution, confirming the high computational cost of futile reasoning.
Table 3: Performance on General Tasks on Qwen3-8B.
MethodAIME 2024GPQA
Acc ↑Reliability ↑Length ↓Acc ↑Reliability ↑Length ↓
Vanilla75.400.754214,78859.850.59857,506
CaRL74.600.785412,41158.330.67685,620
Δ-0.8+3.1-16.1%-1.5+13.1-25.1%
Figure 7: Overview of CaRL. Our framework combines (1) Capability-Calibrated Reward Shaping that establishes a preference hierarchy favoring refusal over hallucination, and (2) Hindsight Refusal Augmentation that converts failed trajectories into refusal trajectories, enabling capability-aligned behavior.
Figure 7: Overview of CaRL. Our framework combines (1) Capability-Calibrated Reward Shaping that establishes a preference hierarchy favoring refusal over hallucination, and (2) Hindsight Refusal Augmentation that converts failed trajectories into refusal trajectories, enabling capability-aligned behavior.
Figure 8: Case study on the countdown task.
Figure 8: Case study on the countdown task.

실제로 확인된 결과

  • 모든 실험 모델(Qwen3-8B, Qwen3-32B, gpt-oss-120b, Qwen3-235B-A22B, DeepSeek-V3.2)이 문제 난이도가 올라가도 거절률이 거의 늘지 않는 '보편적 능력 과신'을 보였고, 가장 어려운 단계(N=8)에서도 명시적 지시 후에도 80% 이상 헛된 시도가 이어졌다.
  • 실패 유형 중 '겉만 번드르르한 추론'이 57~68%로 가장 많았고, '끝없는 시도'는 30~40%로 유지, '반복 루프'는 13%에서 2%로 줄어 난이도가 오를수록 모델이 더 정교하게 틀린 근거를 만들어냄을 확인했다.
  • Qwen3-32B 분석에서 과신(20%)이 과잉신중(3.4%)보다 6배 더 흔했고, 난이도가 오를수록 거절 재현율은 100%에서 30%로 떨어지고 능력 손실은 0%에서 10%로 늘어 자연스러운 프롬프트 유도만으로는 정렬이 어렵다는 것을 확인했다.
  • CaRL 적용 후 Qwen3-8B는 헛된 추론 65.5%→7.0%, 신뢰도 +0.13, Qwen3-14B는 78.6%→1.0%, 신뢰도 +0.16으로 개선됐으며, 아웃오브도메인(Sudoku) 과제와 일반 벤치마크(AIME 2024, GPQA)에서도 정확도 손실이 2% 미만으로 유지됐다.
  • 보상만 조정한 RLunk 변형은 8B 모델에서 헛된 추론 비율이 98~99%로 개선되지 않았고, 정적 데이터로 학습한 RFT는 훈련 분포 밖 과제(Sudoku)에서 성능이 급격히 무너져, CaRL의 두 요소(보상 설계+사후 거절 증강)가 함께 필요함을 보였다.
Figure 9: Initial Reasoning Phase. The model systematically explores combinations (e.g., 97+66=163, 3×51=153), attempting to construct the target value 275. Early attempts show valid mathematical reasoning but fail to reach the exact target.
Figure 9: Initial Reasoning Phase. The model systematically explores combinations (e.g., 97+66=163, 3×51=153), attempting to construct the target value 275. Early attempts show valid mathematical reasoning but fail to reach the exact target.
Figure 10: Final Output After Degenerate Repetition. After 50+ failed attempts, the model outputs (97+66+51+38+37)−(3+3+3)=280 while incorrectly asserting it equals 275. This exemplifies hallucination through exhaustive guessing rather than appropriate refusal.
Figure 10: Final Output After Degenerate Repetition. After 50+ failed attempts, the model outputs (97+66+51+38+37)−(3+3+3)=280 while incorrectly asserting it equals 275. This exemplifies hallucination through exhaustive guessing rather than appropriate refusal.

어디에 쓸 수 있나

  • 수학 퍼즐이나 알고리즘형 추론 과제에서 모델이 스스로 풀 수 없다고 판단될 때 '모르겠다'고 답하게 만드는 훈련 방식을 다른 추론 모델에 적용해 볼 수 있다.
  • 고신뢰성이 요구되는 서비스에서 AI가 틀린 답을 그럴듯하게 제시하는 위험을 줄이기 위한 강화학습 보상 설계 아이디어로 참고할 수 있다.
  • 불필요하게 긴 추론 과정을 조기에 중단시켜 연산 비용을 줄이려는 시스템 설계에 참고할 수 있다.

한계와 남은 검증

  • 실험은 외부 지식이 필요 없는 순수 알고리즘 추론 과제(Countdown, Sudoku)에 한정되어 있어, 지식이 섞인 실제 질의응답이나 수학 문제 등 다른 영역에서 통하는지는 아직 검증되지 않았다.
  • 저자들도 개방형 질의응답이나 수학적 추론 등의 영역으로 CaRL을 확장해 거절 메커니즘이 일반화되는지 검증할 계획이라고 밝혔다.
  • 실험 모델은 Qwen3-8B, Qwen3-14B 두 가지 규모에 한정되어 있어 다른 아키텍처나 훨씬 큰 모델에서도 동일한 효과가 나는지는 확인되지 않았다.

왜 중요한가

그럴듯해 보이지만 틀린 AI의 답변은 사용자가 잘못된 정보를 믿게 만들어 신뢰성이 중요한 분야에서 큰 위험이 된다. 이 연구는 모델이 스스로 한계를 인식하고 '모르겠다'고 말하도록 훈련시키는 구체적 방법을 제시해, AI를 더 안전하게 활용할 수 있는 실질적 단서를 제공한다.

이 논문의 용어

  • 헛된 추론(futile reasoning) · 모델이 자기 능력 밖의 문제를 만났을 때 겉보기엔 정상적이지만 실제로는 쓸모없고 틀린 풀이 과정을 계속 만들어내는 현상
  • CaRL(Capability-aligned Reinforcement Learning) · 모델의 실제 능력 한계에 맞춰 거절 행동을 학습시키는 강화학습 프레임워크
  • 사후 거절 증강(Hindsight Refusal Augmentation) · 실패한 풀이 과정의 마지막 부분만 거절 문구로 바꿔서 부족한 거절 학습 데이터를 늘리는 기법
  • GRPO(Group Relative Policy Optimization) · 여러 응답을 그룹으로 비교해 상대적으로 더 나은 응답을 학습하도록 하는 강화학습 알고리즘
  • Countdown 과제 · 주어진 숫자들과 사칙연산만으로 목표 숫자를 만드는 산술 퍼즐, 게임 24의 변형

저자 · Xinyan Guan

arXiv에서 원문 보기

최신 논문

논문 전체 보기 →

METAL MEDIA 최신 기사

그림 출처: Xinyan Guan et al., arXiv:2607.29211, arxiv-nonexclusive