Persona-Guided LLM Agents for Task-Oriented Dialogue
AI 챗봇이 사용자 성격에 맞춰 말투를 바꾸면 만족도는 오르지만 거짓말도 늘어난다
연구진은 호텔·레스토랑 예약을 도와주는 AI 대화 시스템이 사용자의 성격(외향적, 예민함 등)에 맞춰 반응을 조절할 때 어떤 일이 벌어지는지 실험했다. AI끼리 대화하게 만들어, 시스템이 사용자 성격을 전혀 모를 때, 대화 중 눈치껏 추측할 때, 미리 정답을 알려줄 때 세 조건을 비교했다. 그 결과 성격에 맞춘 대응은 사용자 만족도와 요구사항 충족도를 높였지만, 동시에 사실과 다른 말(허위 정보)을 하는 경우도 늘어났다.
METAL MEDIA 해설 도표
AI 챗봇이 사용자 성격에 맞춰 말투를 바꾸면 만족도는 오르지만 거짓말도 늘어난다
01GPT-4o, Qwen3-Next-80B, Gemini 2.0 Flash 세 개의 대형언어모델을 이용해, 성격이 부여된 '가짜 사용자' AI와 예약을 도와주는 '시스템' AI가 서로 대화하도록 만들었다.
02시스템이 사용자 성격 정보를 전혀 모르는 조건(Neutral), 대화 속 힌트로 추측하는 조건(Try), 성격을 미리 정확히 알고 있는 조건(Oracle) 세 가지로 나눠 비교했다.
03성격을 반영한 두 조건(Try, Oracle) 모두 사용자 요구 충족도와 정보 제공률, 만족도를 높였지만 사실성(허위 정보 방지)은 오히려 떨어뜨렸다.
04사용자가 성격을 뚜렷하게 드러낼수록 Oracle 조건의 만족도 개선 효과가 커졌지만, 대화 힌트로 추측하는 Try 조건은 이런 영향을 덜 받아 더 안정적이었다.
05내향적이거나 폐쇄적인 성격은 AI가 대화 속에서 표현하기 어려워했으며, Qwen3와 Gemini는 까다롭고 적대적인 사용자를 만나면 오히려 똑같이 적대적으로 반응하는 문제를 보였다.
METAL MEDIA이 원문을 바탕으로 재구성한 해설 도표이며, 논문 저자의 원문 figure가 아닙니다.
무엇을 했나
GPT-4o, Qwen3-Next-80B, Gemini 2.0 Flash 세 개의 대형언어모델을 이용해, 성격이 부여된 '가짜 사용자' AI와 예약을 도와주는 '시스템' AI가 서로 대화하도록 만들었다.
시스템이 사용자 성격 정보를 전혀 모르는 조건(Neutral), 대화 속 힌트로 추측하는 조건(Try), 성격을 미리 정확히 알고 있는 조건(Oracle) 세 가지로 나눠 비교했다.
성격을 반영한 두 조건(Try, Oracle) 모두 사용자 요구 충족도와 정보 제공률, 만족도를 높였지만 사실성(허위 정보 방지)은 오히려 떨어뜨렸다.
사용자가 성격을 뚜렷하게 드러낼수록 Oracle 조건의 만족도 개선 효과가 커졌지만, 대화 힌트로 추측하는 Try 조건은 이런 영향을 덜 받아 더 안정적이었다.
내향적이거나 폐쇄적인 성격은 AI가 대화 속에서 표현하기 어려워했으며, Qwen3와 Gemini는 까다롭고 적대적인 사용자를 만나면 오히려 똑같이 적대적으로 반응하는 문제를 보였다.
Table 1: Trait-level task and personality realization outcomes averaged across domains, models, and the three system personality-access conditions. Dialogue Completion and Task Success summarize schema-guided task performance. User Trait Score measures dialogue-level realization of the intended user trait in generated task-oriented dialogues, while BFI Score measures prompt-level semantic alignment with the intended Big Five direction using a separate questionnaire probe.
Trait
Task Outcomes
Personality Realization
Dialogue Completion (%)
Task Success (%)
User Trait Score
BFI Score
Extraversion
79.67
82.12
4.996
5.000
Agreeableness
95.89
87.44
4.900
4.963
Conscientiousness
93.89
83.34
4.313
5.000
Neuroticism
77.22
75.66
4.971
4.967
Openness
85.33
81.12
4.920
4.967
introversive
96.34
91.34
0.653
4.625
disagreeable
78.78
76.22
4.773
5.000
unconscientious
96.34
95.22
3.169
4.778
stable
97.89
88.22
3.473
4.917
closed
98.11
92.00
0.027
4.667
Average
89.94
85.27
3.620
4.888
Table 2: Trait-level system-quality scores and pairwise user-satisfaction preferences averaged across domains and LLMs. N, T, and O denote Neutral, Try, and Oracle, respectively. CS, IR, and TR denote Constraint Satisfaction, Inform Rate, and Truthfulness, each reported on a 1–5 scale, where higher is better. Pairwise satisfaction cells report win-rate percentages for the first condition over the second; for example, O/N = 62/38 indicates that Oracle is preferred in 62% of comparisons and Neutral in 38%. Values may sum to less than 100 due to ties.
Trait
CS
IR
TR
Pairwise Satisfaction
N
T
O
N
T
O
N
T
O
O/N
T/N
T/O
Extraversion
4.313
4.390
4.307
4.573
4.650
4.543
4.647
4.247
3.887
76.00/24.00
64.33/35.00
40.00/58.66
Agreeableness
4.750
4.897
4.840
4.853
4.960
4.937
4.927
4.893
4.830
71.66/27.34
66.34/31.00
41.33/57.00
Conscientiousness
4.653
4.783
4.790
4.780
4.910
4.917
4.937
4.850
4.883
65.34/34.66
59.00/39.67
46.34/53.33
Neuroticism
4.113
4.263
3.920
4.673
4.757
4.677
4.570
4.257
4.170
63.00/36.67
69.34/30.66
62.66/37.00
Openness
4.213
4.353
4.447
4.763
4.803
4.783
4.547
4.223
3.733
85.00/15.00
60.33/39.34
29.66/70.34
introversive
4.753
4.917
4.880
4.857
4.953
4.963
4.930
4.867
4.890
39.33/59.34
71.00/26.00
73.67/25.33
disagreeable
3.797
4.093
4.063
4.283
4.550
4.553
4.653
4.547
4.657
28.33/71.67
47.00/53.00
80.66/19.34
unconscientious
4.720
4.917
4.883
4.800
4.950
4.927
4.880
4.883
4.853
59.00/40.67
66.67/32.66
65.34/33.66
stable
4.827
4.960
4.920
4.903
4.983
4.963
4.900
4.910
4.873
58.34/40.34
67.00/30.66
53.33/45.00
closed
4.877
4.883
4.933
4.927
4.970
4.983
4.910
4.943
4.977
50.33/49.00
63.00/34.00
68.00/30.34
Average
4.502
4.646
4.598
4.741
4.849
4.825
4.790
4.662
4.575
59.63/39.87
63.40/35.20
56.10/43.00
Table 3: Spearman correlations between Avg. User Trait Score and pairwise satisfaction deltas at two levels of analysis. Trait-level uses one point per trait (n=10), averaged across models and domains. Cell-level uses model × domain × trait cells (n=60, Tables 8–9). p∗<.05; p∗∗∗<.001.
Trait-level (n=10)
Cell-level (n=60)
Delta
𝝆
p
𝝆
p
ΔO-N (Oracle − Neutral)
+.685
.029∗
+.421
<.001∗∗∗
ΔT-N (Try − Neutral)
−.212
.556
−.157
.232
Table 4: Task success rates (%) across models, domains, and personality-access conditions. In the Hotel domain, Neutral achieves the highest cross-model average task success, whereas in the Restaurant domain, Oracle performs best, followed by Try.
Domain / Model
Neutral
Try
Oracle
Hotel
Gemini 2.0 Flash
84.40
78.00
77.40
GPT-4o
94.60
93.40
92.80
Qwen3-Next-80B
95.60
96.20
96.40
Cross-model Average
91.53
89.20
88.87
Restaurant
Gemini 2.0 Flash
70.20
73.60
76.00
GPT-4o
80.00
85.20
88.40
Qwen3-Next-80B
84.00
84.20
84.00
Cross-model Average
78.07
81.00
82.80
Table 5: Dialogue-completion rates (%) for GPT-4o, Qwen3, and Gemini across personality traits in the Hotel and Restaurant domains. N, T, and O denote Neutral, Try, and Oracle, respectively. The average row reports the mean dialogue-completion rate across the ten personality traits for each model and domain.
Model
Trait
Hotel (%)
Restaurant (%)
N
T
O
N
T
O
GPT-4o
Extraversion
96.0
90.0
94.0
82.0
98.0
94.0
Agreeableness
100.0
100.0
98.0
82.0
96.0
98.0
Conscientiousness
96.0
94.0
92.0
86.0
96.0
92.0
Neuroticism
84.0
82.0
70.0
80.0
80.0
72.0
Openness
90.0
80.0
90.0
78.0
88.0
90.0
Introversive
96.0
96.0
96.0
88.0
98.0
98.0
disagreeable
84.0
82.0
84.0
56.0
84.0
78.0
unconscientious
100.0
100.0
100.0
74.0
88.0
98.0
stable
100.0
100.0
100.0
94.0
98.0
94.0
closed
98.0
100.0
98.0
98.0
90.0
96.0
Average
94.4
92.4
92.2
81.8
91.6
91.0
Qwen3
Extraversion
70.0
66.0
54.0
80.0
84.0
82.0
Agreeableness
92.0
90.0
92.0
100.0
100.0
100.0
Conscientiousness
96.0
98.0
98.0
100.0
100.0
100.0
Neuroticism
80.0
84.0
76.0
82.0
88.0
84.0
Openness
88.0
84.0
88.0
90.0
80.0
84.0
Introversive
86.0
94.0
98.0
100.0
100.0
100.0
disagreeable
68.0
74.0
74.0
76.0
78.0
86.0
unconscientious
98.0
100.0
100.0
100.0
100.0
100.0
stable
96.0
100.0
96.0
100.0
100.0
100.0
closed
92.0
98.0
100.0
100.0
100.0
100.0
Average
86.6
88.8
87.6
92.8
93.0
93.6
Gemini
Extraversion
50.0
60.0
52.0
96.0
94.0
92.0
Agreeableness
96.0
96.0
92.0
100.0
98.0
96.0
Conscientiousness
70.0
94.0
90.0
96.0
96.0
96.0
Neuroticism
80.0
72.0
66.0
68.0
66.0
76.0
Openness
80.0
84.0
74.0
94.0
82.0
92.0
Introversive
98.0
100.0
88.0
100.0
100.0
98.0
disagreeable
66.0
70.0
76.0
90.0
92.0
100.0
Table 6: Restaurant-domain system-evaluation results for GPT-4o, Qwen3, and Gemini across personality traits under Neutral (N), Try (T), and Oracle (O) conditions. The table reports Constraint Satisfaction, Truthfulness, and Inform Rate; the average row reports the mean across the ten personality traits for each model.
Model
Trait
Constraint Satisfaction
Truthfulness
Inform Rate
N
T
O
N
T
O
N
T
O
GPT-4o
Extraversion
4.22
4.88
4.94
4.90
4.66
4.78
4.38
5.00
4.98
Agreeableness
4.32
4.84
4.90
4.80
4.58
4.64
4.38
4.84
5.00
Conscientiousness
4.60
4.94
4.76
4.86
4.82
4.82
4.76
4.98
4.84
Neuroticism
4.30
4.40
4.04
4.12
4.02
3.84
4.62
4.68
4.48
Openness
4.26
4.54
4.60
4.64
4.70
4.66
4.44
4.58
4.66
Introversive
4.48
4.90
4.84
4.98
4.76
4.90
4.62
4.92
4.98
disagreeable
3.16
4.12
4.40
4.50
4.28
4.30
3.32
4.56
4.58
unconscientious
4.04
4.76
4.86
4.78
4.92
4.70
4.06
4.76
4.96
stable
4.84
5.00
4.88
4.76
4.76
4.66
4.80
5.00
4.94
closed
4.80
4.68
4.72
5.00
5.00
4.94
4.92
4.88
4.90
Average
4.30
4.71
4.69
4.73
4.65
4.62
4.43
4.82
4.83
Qwen3
Extraversion
4.68
4.70
4.62
4.12
3.72
3.66
4.76
4.78
4.70
Agreeableness
4.96
5.00
5.00
5.00
5.00
4.92
4.96
5.00
5.00
Conscientiousness
5.00
5.00
4.96
4.96
5.00
4.94
5.00
5.00
4.96
Neuroticism
3.92
4.22
3.92
4.74
4.52
4.86
4.68
4.78
4.86
Openness
4.10
4.44
4.52
4.70
4.26
3.40
4.78
4.82
4.76
Introversive
5.00
5.00
5.00
4.94
4.94
4.96
5.00
5.00
5.00
disagreeable
2.78
3.18
3.14
4.38
4.72
4.74
4.16
4.38
4.36
unconscientious
5.00
5.00
5.00
5.00
4.98
5.00
5.00
5.00
5.00
stable
4.94
5.00
4.98
5.00
5.00
5.00
5.00
5.00
5.00
closed
5.00
5.00
5.00
5.00
5.00
5.00
5.00
5.00
5.00
Average
4.54
4.65
4.61
4.78
4.71
4.65
4.83
4.88
4.86
Gemini
Extraversion
4.92
4.94
4.70
4.94
4.68
4.20
4.96
4.96
4.76
Agreeableness
5.00
4.92
4.88
4.96
4.88
4.98
5.00
4.92
4.90
Conscientiousness
4.86
4.92
4.96
5.00
4.64
4.86
4.96
4.96
5.00
Neuroticism
4.08
4.14
4.20
4.92
4.70
4.42
4.62
4.78
4.82
Openness
4.76
4.54
4.72
4.84
4.48
4.48
4.90
4.88
4.86
Introversive
4.98
5.00
4.96
5.00
5.00
4.92
5.00
5.00
5.00
disagreeable
4.64
4.84
4.80
4.96
4.80
4.98
4.66
4.90
5.00
Table 7: Hotel-domain system-evaluation results for GPT-4o, Qwen3, and Gemini across personality traits under Neutral (N), Try (T), and Oracle (O) conditions. The table reports Constraint Satisfaction, Truthfulness, and Inform Rate; the average row reports the mean across the ten personality traits for each model.
Model
Trait
Constraint Satisfaction
Truthfulness
Inform Rate
N
T
O
N
T
O
N
T
O
GPT-4o
Extraversion
4.92
4.76
4.74
4.74
4.36
4.14
4.96
4.82
4.80
Agreeableness
5.00
5.00
5.00
4.96
5.00
4.84
5.00
5.00
5.00
Conscientiousness
4.82
4.68
4.82
4.94
4.92
4.86
4.94
4.84
4.96
Neuroticism
4.42
4.58
4.10
4.62
4.24
4.24
4.84
4.80
4.80
Openness
4.30
4.24
4.58
4.48
3.74
3.36
4.74
4.76
4.88
Introversive
4.82
4.88
4.92
5.00
4.82
4.82
4.90
4.92
5.00
disagreeable
4.68
4.60
4.40
4.82
4.90
4.82
4.86
4.74
4.72
unconscientious
5.00
5.00
4.96
5.00
5.00
4.90
5.00
5.00
5.00
stable
4.96
5.00
5.00
4.94
5.00
4.94
5.00
5.00
5.00
closed
5.00
4.96
4.92
5.00
5.00
5.00
5.00
5.00
5.00
Average
4.79
4.77
4.74
4.85
4.70
4.59
4.92
4.89
4.92
Qwen3
Extraversion
3.76
3.58
3.42
4.18
3.66
2.56
4.68
4.58
4.24
Agreeableness
4.48
4.70
4.52
4.84
4.98
4.82
4.94
5.00
4.88
Conscientiousness
4.66
4.46
4.56
4.94
4.82
4.96
4.96
4.90
4.94
Neuroticism
3.40
3.68
3.22
4.40
3.64
3.26
4.54
4.62
4.42
Openness
3.30
3.70
4.16
3.76
3.82
2.70
4.90
4.86
4.90
Introversive
4.32
4.72
4.88
4.72
4.72
4.88
4.72
4.88
4.96
disagreeable
3.66
3.80
3.52
4.40
3.98
4.28
4.28
4.32
4.34
unconscientious
4.64
4.78
4.84
4.78
4.82
4.82
4.96
4.96
4.94
stable
4.56
4.76
4.84
4.82
4.82
4.84
4.88
4.96
4.92
closed
4.62
4.86
5.00
4.62
4.80
4.98
4.78
4.94
5.00
Average
4.14
4.30
4.30
4.55
4.41
4.21
4.76
4.80
4.75
Gemini
Extraversion
3.38
3.48
3.42
5.00
4.40
3.98
3.70
3.76
3.78
Agreeableness
4.74
4.92
4.74
5.00
4.92
4.78
4.84
5.00
4.84
Conscientiousness
3.98
4.70
4.68
4.92
4.90
4.86
4.06
4.78
4.80
Neuroticism
4.56
4.56
4.04
4.62
4.42
4.40
4.74
4.88
4.68
Openness
4.56
4.66
4.10
4.86
4.34
3.80
4.82
4.92
4.64
Introversive
4.92
5.00
4.68
4.94
4.96
4.86
4.90
5.00
4.84
disagreeable
3.86
4.02
4.12
4.86
4.60
4.82
4.42
4.40
4.32
Table 8: Hotel-domain pairwise user-satisfaction win rates with averaged personality-realization metrics for GPT-4o, Qwen3, and Gemini. O, T, and N denote Oracle, Try, and Neutral, respectively. Each comparison cell reports the win rate of the first condition over the second condition; for example, O/N = 36/58 means Oracle is preferred in 36% of paired comparisons and Neutral in 58%. Percentages may sum to less than 100 when ties occur. BFI Score is the intended-direction score from the BFI probe for the target trait. Avg. User Trait Score and Avg. Rank are averaged across Oracle, Try, and Neutral cases. Lower Avg. Rank indicates stronger realization of the intended trait.
Model
Trait
O / N (%)
T / N (%)
T / O (%)
BFI Score
Avg. User Trait Score
Avg. Rank
GPT-4o
introversive
36/58
64/26
70/24
4.500
0.207
3.913
closed
28/70
50/40
72/26
4.100
0.020
3.987
stable
44/56
60/34
62/36
4.750
3.173
2.653
Extraversion
80/20
64/36
34/62
5.000
4.993
1.000
Agreeableness
76/20
56/36
44/52
4.889
4.887
1.000
Conscientiousness
46/54
44/54
42/56
5.000
4.047
1.327
Neuroticism
54/44
60/40
58/40
5.000
4.920
1.007
Openness
84/16
54/46
28/72
5.000
4.793
1.013
disagreeable
44/56
38/62
66/34
5.000
4.580
1.020
unconscientious
44/54
54/44
64/30
4.444
2.940
2.187
Average
53.6/44.8
54.4/41.8
54.0/43.2
4.768
3.456
1.911
Qwen3
introversive
52/48
56/40
44/56
4.500
0.267
3.873
closed
60/38
56/44
44/56
5.000
0.000
4.000
stable
66/34
54/44
34/66
5.000
3.553
1.833
Extraversion
60/40
58/42
54/46
5.000
5.000
1.000
Agreeableness
76/24
68/32
30/70
5.000
4.820
1.000
Conscientiousness
72/28
48/52
32/68
5.000
4.053
1.053
Neuroticism
70/30
74/26
60/40
5.000
4.993
1.000
Openness
88/12
50/50
16/84
5.000
4.987
1.000
disagreeable
22/78
34/66
74/26
5.000
4.780
1.020
unconscientious
60/40
68/30
54/46
5.000
2.193
2.553
Average
62.6/37.2
56.6/42.6
44.2/55.8
4.950
3.465
1.833
Gemini
introversive
30/70
70/30
86/14
4.875
1.487
3.133
closed
54/46
64/36
74/26
4.900
0.060
3.980
stable
48/50
76/24
64/36
5.000
3.693
1.753
Extraversion
68/32
68/32
44/54
5.000
4.993
1.000
Agreeableness
76/24
70/30
44/56
5.000
4.993
1.000
Conscientiousness
64/36
66/34
64/36
5.000
4.840
1.000
Neuroticism
70/30
84/16
78/22
5.000
5.000
1.000
Openness
92/8
92/8
34/66
4.900
4.980
1.007
Table 9: Restaurant-domain pairwise user-satisfaction win rates with averaged personality-realization metrics for GPT-4o, Qwen3, and Gemini. O, T, and N denote Oracle, Try, and Neutral, respectively. Each comparison cell reports the win rate of the first condition over the second condition, e.g., O/N = 62/38 means Oracle is preferred in 62% of paired comparisons and Neutral in 38%. Percentages may sum to less than 100 when ties occur. BFI Score is the intended-direction score from the BFI probe for the target trait. Avg. User Trait Score and Avg. Rank are averaged across Oracle, Try, and Neutral cases. Lower Avg. Rank indicates stronger realization of the intended trait.
Model
Trait
O / N (%)
T / N (%)
T / O (%)
BFI Score
Avg. User Trait Score
Avg. Rank
GPT-4o
introversive
62/38
82/18
60/40
4.500
1.420
3.173
closed
54/46
58/36
60/34
4.100
0.393
3.760
stable
66/34
76/24
62/36
4.750
3.260
2.213
Extraversion
88/12
66/32
28/70
5.000
4.973
1.000
Agreeableness
76/24
78/18
38/56
4.889
4.593
1.020
Conscientiousness
70/30
70/30
54/46
5.000
3.507
1.807
Neuroticism
50/50
58/42
62/38
5.000
4.673
1.020
Openness
80/20
60/38
40/60
5.000
4.827
1.020
disagreeable
70/30
64/36
64/36
5.000
4.527
1.000
unconscientious
70/30
78/22
68/32
4.444
1.260
3.413
Average
68.6/31.4
69.0/29.6
53.6/44.8
4.768
3.343
1.943
Qwen3
introversive
4/96
70/26
98/2
4.500
0.240
3.867
closed
42/58
72/26
84/16
5.000
0.000
4.000
stable
76/22
70/26
38/60
5.000
3.580
1.747
Extraversion
72/28
62/36
48/52
5.000
4.987
1.000
Agreeableness
88/10
68/28
18/82
5.000
4.787
1.013
Conscientiousness
74/26
76/18
38/62
5.000
4.073
1.020
Neuroticism
50/50
68/32
70/30
5.000
4.980
1.000
Openness
82/18
52/48
20/80
5.000
4.913
1.040
disagreeable
20/80
58/42
86/14
5.000
4.960
1.000
unconscientious
60/40
70/30
72/28
5.000
2.027
2.660
Average
56.8/42.8
66.6/31.2
57.2/42.6
4.950
3.455
1.835
Gemini
introversive
52/46
84/16
84/16
4.875
1.895
2.969
closed
64/36
78/22
74/24
4.900
0.140
3.940
stable
50/46
66/32
60/36
5.000
3.660
1.820
Extraversion
88/12
68/32
32/68
5.000
5.000
1.000
Agreeableness
38/62
58/42
74/26
5.000
5.000
1.000
Conscientiousness
66/34
50/50
48/52
5.000
4.600
1.000
Neuroticism
84/16
72/28
48/52
5.000
5.000
1.000
Openness
84/16
54/46
40/60
4.900
4.953
1.087
Table 10: Per-trait Avg. Intended Score and average pairwise satisfaction deltas (pp), averaged across six model–domain combinations. Traits are sorted by Avg. Intended Score in descending order. The disagreeable trait is discussed separately in Appendix C.3.
Trait
Avg. Intended Score
Avg. ΔO−N (%)
Avg. ΔT−N (%)
Extraversion
4.991
+52.0
+29.3
Neuroticism
4.928
+26.3
+38.7
Openness
4.909
+70.0
+21.0
Agreeableness
4.847
+44.3
+35.3
disagreeable†
4.788
−43.3
−6.0
Conscientiousness
4.187
+30.7
+19.3
stable
3.486
+18.0
+36.3
unconscientious
2.737
+18.3
+34.0
introversive
0.919
−20.0
+45.0
closed
0.102
+1.3
+29.0
왜 중요한가
고객센터나 예약 서비스처럼 AI가 사람과 여러 차례 대화해야 하는 서비스를 만들 때, 사용자 성향에 맞춰 대응을 바꾸는 것이 만족도를 높이지만 정확성을 해칠 수 있다는 트레이드오프를 구체적 수치로 보여준다. 추가 학습 없이 프롬프트만으로 이런 성격 적응형 대화 시스템을 만들 수 있다는 점도 실무적으로 유용하다.
이 논문의 용어
Big Five 성격 특성 · 외향성, 친화성, 성실성, 신경성, 개방성 등 성격을 다섯 축으로 나누는 심리학 모델
TOD(과업지향 대화) · 예약이나 검색처럼 구체적 목표를 완수하기 위해 여러 차례 주고받는 대화
Constraint Satisfaction · 시스템의 최종 추천이나 예약이 사용자가 요구한 조건을 실제로 충족했는지 평가하는 지표
Truthfulness(사실성) · 시스템이 근거 없는 내용을 지어내지 않고 실제 대화·검색 결과에 기반해 답했는지 평가하는 지표
Spearman 상관계수 · 두 값이 함께 증가·감소하는 경향을 순위 기반으로 측정하는 통계 지표
본문에 싣지 못한 그림
Figure 1: Overview of our framework for personality-aware task-oriented dialogue. A personality-conditioned user agent interacts with a system agent under three personality-access conditions—Neutral, Try, and Oracle—within an SGD-based API-grounded task setting. Generated dialogues are evaluated for task outcomes, system quality, and personality-related behavior.
Figure 2: Spearman correlation between Avg. User Intended Score and pairwise satisfaction deltas across traits (n=10). Oracle gains increase with trait realization (ρ=+0.685, p=.029), whereas Try gains do not (ρ=−0.103, p=.777). Point size indicates cross-dataset SD; disagreeable reflects a Qwen3/Gemini mirroring artifact (Appendix C.3).
Figure 3: Representative GPT-4o Restaurant-domain score-delta figures. The top panel shows score changes from Neutral to Try, and the bottom panel shows score changes from Neutral to Oracle.