Hear2Act: Benchmarking When Prosody Should Change What an Assistant Does
말투(억양)가 정말 AI 비서의 판단을 바꾸는가를 480개 시나리오로 시험한 벤치마크
이 논문은 사용자가 같은 말을 해도 목소리 톤이 다르면 AI 비서가 다르게 행동해야 하는 상황을 480개로 만들어 'Hear2Act'라는 평가 틀을 제안했다. 목소리 정보를 그냥 들려주기만 하면 오디오를 처리할 수 있는 AI 모델도 별 차이를 못 만들었지만, 목소리에서 읽어낸 감정 상태를 텍스트로 명시적으로 적어주면 성능이 크게 올랐다. 즉 현재의 음성 AI는 말투에서 정보를 어느 정도 알아채긴 하지만, 그걸 행동으로 제대로 옮기지는 못한다는 것을 보여준다.
METAL MEDIA 해설 도표
말투(억양)가 정말 AI 비서의 판단을 바꾸는가를 480개 시나리오로 시험한 벤치마크
01연구팀은 항공권 예약, 주거, 금융 등 48개 생활 서비스 영역에서 480개 시나리오를 만들었다. 각 시나리오에는 사용자가 겉으로 말하지 않는 3단계 숨은 고민(반드시 지켜야 할 조건, 강한 선호, 약한 선호)이 있고, 11개의 선택지 중 이 숨은 고민을 모두 만족하는 정답이 정해져 있다.
02같은 시나리오를 두 가지 방식으로 실험했다. 하나는 사용자의 고민을 말로 명확히 표현하는 경우(Explicit lexical), 다른 하나는 말은 '괜찮아요' 같은 수긍이지만 목소리 톤에서만 망설임이 드러나는 경우(Prosody-mediated)다.
03목소리 톤으로만 신호가 오는 상황에서, 오디오 처리가 가능한 AI 두 종에게 텍스트 대본만 줬을 때 정답률은 14.6%였고, 오디오까지 함께 줘도 15.3%로 거의 그대로였다. 반면 AI가 오디오에서 사용자 상태를 스스로 판단해 글로 적어놓고 그걸 참고해 다음 행동을 정하게 하자 정답률이 39.6%까지 올라 실제 정답 상태를 알려줬을 때(40.7%)에 근접했다.
04이 차이는 사용자의 고민이 이미 말로 명확히 표현된 상황에서는 거의 사라졌다. 즉 목소리 톤이 중요해지는 건 말로 다 전달되지 않는 정보가 있을 때뿐이었고, 정답 정보를 정확히 알려주자 AI는 성급하게 결론짓지 않고 더 캐물으며(질문 비율 증가) 사용자의 숨은 요구를 더 많이 끌어냈다(모든 고민을 다 알아낸 비율이 3%에서 30%로 상승).
05사람 청취자들에게 합성 음성을 들려준 결과 92% 정확도로 목소리 톤의 의도(해결됐는지 안 됐는지)를 구분해냈고, 이는 음성 합성 자체는 신뢰할 만하다는 것을 보여준다. 반면 AI 모델은 목소리에서 '문제가 남아있음'을 읽어내는 정확도가 사람보다 낮아, 특히 아직 해결 안 된 신호를 놓치는 경우가 많았다.
METAL MEDIA이 원문을 바탕으로 재구성한 해설 도표이며, 논문 저자의 원문 figure가 아닙니다.
무엇을 했나
연구팀은 항공권 예약, 주거, 금융 등 48개 생활 서비스 영역에서 480개 시나리오를 만들었다. 각 시나리오에는 사용자가 겉으로 말하지 않는 3단계 숨은 고민(반드시 지켜야 할 조건, 강한 선호, 약한 선호)이 있고, 11개의 선택지 중 이 숨은 고민을 모두 만족하는 정답이 정해져 있다.
같은 시나리오를 두 가지 방식으로 실험했다. 하나는 사용자의 고민을 말로 명확히 표현하는 경우(Explicit lexical), 다른 하나는 말은 '괜찮아요' 같은 수긍이지만 목소리 톤에서만 망설임이 드러나는 경우(Prosody-mediated)다.
목소리 톤으로만 신호가 오는 상황에서, 오디오 처리가 가능한 AI 두 종에게 텍스트 대본만 줬을 때 정답률은 14.6%였고, 오디오까지 함께 줘도 15.3%로 거의 그대로였다. 반면 AI가 오디오에서 사용자 상태를 스스로 판단해 글로 적어놓고 그걸 참고해 다음 행동을 정하게 하자 정답률이 39.6%까지 올라 실제 정답 상태를 알려줬을 때(40.7%)에 근접했다.
이 차이는 사용자의 고민이 이미 말로 명확히 표현된 상황에서는 거의 사라졌다. 즉 목소리 톤이 중요해지는 건 말로 다 전달되지 않는 정보가 있을 때뿐이었고, 정답 정보를 정확히 알려주자 AI는 성급하게 결론짓지 않고 더 캐물으며(질문 비율 증가) 사용자의 숨은 요구를 더 많이 끌어냈다(모든 고민을 다 알아낸 비율이 3%에서 30%로 상승).
사람 청취자들에게 합성 음성을 들려준 결과 92% 정확도로 목소리 톤의 의도(해결됐는지 안 됐는지)를 구분해냈고, 이는 음성 합성 자체는 신뢰할 만하다는 것을 보여준다. 반면 AI 모델은 목소리에서 '문제가 남아있음'을 읽어내는 정확도가 사람보다 낮아, 특히 아직 해결 안 된 신호를 놓치는 경우가 많았다.
Figure 1: Hear2Act overview. (1) Each scenario combines a surface request, prioritized hidden concerns, and candidates with verifiable satisfaction signatures. (2) The assistant interacts under different feedback and access conditions. (3) Matched rollouts are compared on task outcomes and interaction behavior.
Table 1: Benchmark positioning. P: prosodic input, C: matched control of prosodic access with fixed lexical content, M: multi-turn task decisions, N: task-grounded hidden user need, and O: verifiable trajectory and outcome. △ marks structured user goals conveyed lexically rather than hidden needs.
Benchmark
P
C
M
N
O
MultiWOZ, SGD (3; 19)
×
×
✓
△
✓
SpokenWOZ (21)
✓
×
✓
△
✓
StyleTalk, ParaS2S (16; 29)
✓
✓
×
×
×
MULTI-Bench, HumDial-EIBench (8; 24)
✓
×
✓
×
×
Hear2Act (ours)
✓
✓
✓
✓
✓
Figure 2: Illustrative Hear2Act trajectory under Prosody-mediated feedback. Three representative candidates are shown from the full 11-candidate set. Transcript-only access may confirm prematurely, while ground-truth concern-state access supports further elicitation and selection of the best-fitting option.
Table 2: Hear2Act benchmark and rollout coverage. The 480 model-independent scenarios expand to 54,240 evaluation rollouts across models, access conditions, renderers, and interventions.
Benchmark artifact
SGD-seeded domains
48
Scenarios per domain
10
Benchmark scenarios
480
Candidate options per scenario
11
Hidden concern layers per scenario
3
Feedback realizations
2
Base episode specifications
960
Assistant turn budget
20
Evaluation rollouts
Text LLM main grid
19,200
Text label interventions
1,440
Spoken assistant with Qwen3-TTS
6,720
Spoken assistant with VoxCPM2
6,720
Qwen2-Audio, three rollouts per scenario
20,160
Total evaluation rollouts
54,240
Figure 3: Change in assistant action composition with concern-state access. Points show the T+S−T change, in percentage points, in the share of recommend, ask, and clarify decision turns under Prosody-mediated (red circles) and Explicit lexical (blue triangles) feedback; means are macro-averaged across models.
Table 3: Spoken-assistant results under Prosody-mediated feedback with Qwen3-TTS. T, A, S, and S^ denote transcript, audio, ground-truth state, and audio-inferred state; audio-derived representations are textualized and paired with the transcript. Average is computed across the two audio-capable LLMs. Bold/underline indicate the best/second-best value per column. See Table 4 for Explicit-lexical results and Appendix B for VoxCPM2.
Qwen2.5-Omni
Qwen2-Audio
Average
Input / representation
1st %↑
OptSat %↑
SvcSat %↑
1st %↑
OptSat %↑
SvcSat %↑
1st %↑
OptSat %↑
SvcSat %↑
Direct input
Transcript only (T)
15.4
22.9
12.3
13.7
20.1
12.0
14.6
21.5
12.2
Audio only (A)
17.3
22.7
13.3
14.7
22.4
12.2
16.0
22.6
12.8
Audio + transcript (A+T)
15.9
24.0
14.6
14.7
23.3
13.3
15.3
23.7
14.0
Textualized prosodic representations
Generic affect (HuBERT)
35.7
40.3
21.4
29.4
31.3
19.9
32.6
35.8
20.7
Generic affect (SpeechBrain)
32.6
38.7
20.6
25.7
29.1
18.3
29.2
33.9
19.5
Task-aligned state (T+S^)
43.0
41.8
22.7
36.2
33.4
20.6
39.6
37.6
21.7
Ground-truth state (T+S)
44.5
41.8
24.4
36.9
37.5
23.5
40.7
39.7
24.0
Figure 4: Outcome tracks label fidelity, not asking frequency. Optimal-solution rate on the 48-scenario intervention subset under Prosody-mediated feedback, pooled over five text LLMs. All label conditions have similar ask shares (44–45%).
Table 4: Spoken-assistant use of concern information under Explicit lexical feedback with Qwen3-TTS (Qwen2.5-Omni n=480, Qwen2-Audio n=1,440 per condition). Notation follows Table 3. Results are similar because the concern is explicit in the transcript. Bold/underline mark column-wise highest/next-highest values.
Qwen2.5-Omni
Qwen2-Audio
Average
Input / representation
1st %↑
OptSat %↑
SvcSat %↑
1st %↑
OptSat %↑
SvcSat %↑
1st %↑
OptSat %↑
SvcSat %↑
Direct input
Transcript only (T)
56.2
45.3
26.9
48.3
42.1
26.3
52.3
43.7
26.6
Audio only (A)
55.3
45.7
26.4
48.0
41.3
26.0
51.7
43.5
26.2
Audio + transcript (A+T)
52.2
45.4
26.5
47.6
41.0
25.9
49.9
43.2
26.2
Textualized prosodic representations
Generic affect (HuBERT)
52.8
46.0
25.3
48.5
40.6
24.5
50.7
43.3
24.9
Generic affect (SpeechBrain)
52.8
45.9
25.3
45.0
39.9
23.7
48.9
42.9
24.5
Task-aligned state (T+S^)
52.4
45.8
25.4
49.8
41.2
25.1
51.1
43.5
25.3
Ground-truth state (T+S)
49.7
46.6
26.6
50.3
42.3
26.1
50.0
44.5
26.4
Figure 5: Concern-state access increases hidden-concern disclosure. Bars show mean revealed concerns, dots show per-model means, and in-bar percentages show rollouts revealing all three (T: transcript only; T+S: transcript plus turn-level ground-truth concern-state tags). The effect is concentrated under Prosody-mediated feedback, where concerns remain lexically implicit.
Table 5: Effect of ground-truth concern-state access on text LLMs. Results use 480 matched scenarios with two runs each (n=960 per cell). T is transcript only; T+S adds turn-level ground-truth state tags. Bold marks the larger value in each pair; action composition appears in Figure 3.
Prosody-mediated feedback
Explicit lexical feedback
1st%↑
OptSat%↑
SvcSat%↑
1st%↑
OptSat%↑
SvcSat%↑
Model
T
T+S
T
T+S
T
T+S
T
T+S
T
T+S
T
T+S
Claude Opus 4.6
49.7
79.1
37.0
41.1
24.7
26.2
74.9
81.9
44.6
43.0
25.9
26.4
Kimi K2.5
27.9
69.2
27.7
36.1
14.4
19.6
69.4
74.9
37.1
36.2
17.7
17.7
GLM-5
22.4
60.1
26.1
33.4
13.0
15.8
67.3
73.5
36.8
35.7
15.5
16.0
Qwen3-32B
21.1
53.1
31.0
35.5
10.4
13.7
64.3
64.8
43.1
37.7
12.9
12.6
DeepSeek-V3.2
13.3
56.2
16.4
28.8
6.5
14.4
54.8
67.6
31.8
32.6
14.3
15.1
Five-model mean
26.9
63.5
27.6
35.0
13.8
17.9
66.1
72.5
38.7
37.0
17.3
17.6
Table 6: Domain inventory. Hear2Act covers 48 consumer-service domains grouped into eight families. Each domain is expanded into ten scenario seeds, yielding 480 benchmark scenarios.
Home cleaning services, home renovation contractors, home security systems, landscaping contractors, lawn care contractors, solar panel installation, kitchen appliance upgrades, and mattress replacement.
Finance & Insurance
9
Credit card applications, mortgage lender comparison, investment portfolio allocation, retirement planning advisors, tax preparation services, car insurance policies, health insurance plans, pet insurance policies, and business insurance coverage.
Health & Wellness
7
Dermatologist appointments, pediatrician selection, mental health therapists, meditation apps, gym membership options, fitness tracker devices, and prescription eyeglasses.
Education & Career
5
College major selection, online coding bootcamps, language learning platforms, professional development courses, and laptop purchase for students.
Family & Lifestyle
6
Children’s daycare centers, dog training classes, online dating platforms, wine club memberships, meal delivery subscriptions, and video game purchases.
Media & Devices
4
Cable TV packages, streaming service subscriptions, podcast hosting services, and smartphone upgrades.
Professional Services
4
Legal consultation services, auto mechanic services, business accounting software, and freelance graphic designers.
Total
48
480 scenarios across ten seeds per domain
Table 7: One expanded domain. Ten subdomain seeds for budget airline tickets illustrate variation in user situation and decision pressure before scenario instantiation.
Subdomain situation (who / pressure)
Opening request
1
Last-minute emergency travel for family medical situation with extremely limited budget
“I need to fly to see my sick grandmother tomorrow but only have $200—what are my cheapest options?”
2
College student planning spring break trip with friends on tight budget
“Can you help me find the cheapest flights for four college students going to Miami for spring break?”
3
Budget-conscious family of five planning annual vacation
“What’s the most affordable way to fly my family of five to Orlando for our Disney World trip?”
4
Digital nomad seeking flexible travel dates for extended European backpacking
“I want to backpack through Europe for 3 months—which budget airlines offer the best multi-city deals?”
5
Job interview candidate needing quick affordable travel for an unexpected opportunity
“I have a job interview in Seattle next week and need the cheapest flight possible from Chicago.”
6
Retiree on fixed income wanting to visit grandchildren regularly
“As a senior on a fixed income, what budget airline options exist for regular visits to see my grandkids?”
7
Young professional attending a destination wedding with multiple flight segments
“I need budget flights to get to my friend’s wedding in Bali, including connections—what’s the cheapest route?”
8
Small business owner traveling frequently for client meetings on a startup budget
“I need to travel monthly for business but my startup has a tight travel budget—which airlines offer the best deals for frequent short trips?”
9
International student trying to visit home during semester break
“I’m an international student wanting to fly home to India for winter break—what are the most affordable long-haul options?”
10
Adventure traveler planning a multi-stop trip to remote destinations
“I want to visit three different countries in South America on a backpacker’s budget—which budget airlines serve those routes?”
Table 8: Full spoken-assistant condition grid for Qwen2.5-Omni with VoxCPM2 (n=480 per condition). T, A, S, and S^ denote transcript, audio, ground-truth concern state, and audio-inferred concern state. Audio-derived representations are supplied as text alongside the transcript. Bold/underline mark the highest/next-highest value per column.
Prosody-mediated feedback
Explicit lexical feedback
Input / representation
1st %↑
OptSat %↑
SvcSat %↑
1st %↑
OptSat %↑
SvcSat %↑
Direct input
Transcript only (T)
12.5
22.5
11.9
53.2
44.4
28.8
Audio only (A)
16.1
22.6
14.1
53.9
45.4
28.2
Audio + transcript (A+T)
16.9
25.6
14.8
48.4
44.9
28.3
Textualized prosodic representations
Generic affect (HuBERT)
35.5
40.0
23.9
52.0
44.5
27.6
Generic affect (SpeechBrain)
32.4
39.8
24.1
51.8
44.6
27.6
Task-aligned concern state (T+S^)
42.6
42.0
25.5
51.4
44.9
28.5
Ground-truth concern state (T+S)
45.7
42.5
27.6
52.2
45.5
30.0
Table 9: Scenario-level paired bootstrap 95% CIs for key diagnostic contrasts (2,000 joint scenario resamples). Point estimates correspond to Tables 5, 3, and 4, and Figure 4. T, A, S, and S^ denote transcript, audio, ground-truth concern state, and audio-inferred textual state, respectively. The label-fidelity block uses 48 intervention scenarios.
Contrast
Prosody-mediated
Explicit lexical
Text LLMs, pooled over five models (Δ=+State−Base)
1st%
+36.7 [+34.3, +38.9]
+6.4 [+4.7, +8.1]
OptSat%
+7.3 [+5.8, +8.7]
−1.6 [−2.6, −0.8]
SvcSat%
+4.1 [+2.8, +5.4]
+0.3 [−0.3, +1.0]
Qwen2.5-Omni-7B, 1st% diagnostic contrasts
Ground-truth state on transcript ((T+S)−T)
+29.0 [+23.6, +34.4]
−6.5 [−12.9, +0.0]
Audio-inferred vs. ground-truth state ((T+S^)−(T+S))
Audio-inferred vs. ground-truth state ((T+S^)−(T+S))
−0.7 [−4.3, +3.0]
−0.6 [−4.2, +3.1]
Label-fidelity ladder, 1st% successive steps (48 scenarios, pooled models)
All-positive − no state
+0.6 [−3.8, +5.4]
−5.0 [−11.0, +0.4]
Shuffled − all-positive
+7.1 [+0.8, +12.9]
+3.3 [−2.9, +9.6]
All-negative − shuffled
+10.8 [+3.8, +17.5]
+2.1 [−3.3, +7.5]
Correct state − all-negative
+14.8 [+8.3, +21.7]
+5.2 [+0.0, +10.8]
Table 10: Speech-rendering validation. Human listeners recover resolved versus unresolved concern status from both renderers with 0.92 accuracy (n=100 per renderer; balanced classes), confirming that the intended prosodic contrast remains perceptible after rendering. κ denotes inter-annotator agreement.
Concern status
Qwen3-TTS
VoxCPM2
Concern resolved (O+)
0.93
0.90
Concern unresolved (O−)
0.91
0.93
Overall
0.92
0.92
κ (annotators)
0.96
0.90
Table 11: Qwen2.5-Omni-7B as a concern-cue reader: accuracy against the intended concern status on the audited clips (100 per renderer, balanced 50/50; protocol of Section 4.4). Human values average the two annotators. Bottom block: inter-annotator κ; raw model–annotator agreement; model–annotator κ (all averaged over the two annotators).
Qwen3-TTS
VoxCPM2
Human
Model
Human
Model
Resolved (O+)
0.93
0.86
0.90
0.88
Unresolved (O−)
0.91
0.84
0.93
0.64
Overall
0.92
0.85
0.92
0.76
κ, annotators
0.96
0.90
Agreement, model
0.87
0.78
κ, model
0.74
0.55
Table 12: State-to-delivery mapping. Each concern state is mapped to graded delivery labels used for speech realization.
Concern state
Delivery labels
Resolved: genuine acceptance
satisfied, warm, enthusiastic, relieved
Unresolved: reluctant acceptance
underwhelmed, lukewarm, hesitant, flat
Unresolved: voiced concern
concerned
Unresolved: rejection
frustrated, disappointed, impatient, firm
왜 중요한가
음성 비서나 콜센터 AI가 사용자의 말투에서 불만이나 망설임을 읽어내지 못하면, 겉으로는 수긍한 것처럼 보이는 상황에서 잘못된 결정을 내려 실제 문제를 해결하지 못할 수 있다. 이 연구는 오디오를 처리하는 AI라도 목소리 정보를 명시적인 텍스트로 바꿔주지 않으면 실제 판단에 잘 반영하지 못한다는 점을 보여줘, 음성 AI 설계 시 '중간 표현'을 만드는 과정이 왜 필요한지 근거를 제시한다.
이 논문의 용어
프로소디(prosody) · 말의 높낮이, 억양, 속도, 톤 등 단어 자체가 아닌 목소리의 전달 방식
과제지향 대화(task-oriented dialogue) · 항공권 예약처럼 구체적인 목표 달성을 위해 여러 차례 주고받는 대화
오디오 처리 가능 LLM(audio-capable LLM) · 텍스트뿐 아니라 음성 파일을 직접 입력받아 이해할 수 있는 대형 언어모델
최적해율(optimal-solution rate) · 전체 대화 중 진짜 최선의 선택지로 끝난 비율
숨은 고민(hidden concern) · 사용자가 먼저 말하지 않지만 실제로는 중요하게 여기는 요구사항
저자 · Xinyi Liu, Hooshang Nayyeri, Dilek Hakkani-Tur, Emine Yilmaz, JK Kim, Yifei Zhang, Charith Peris, Hari Thadakamalla