Everything K-culture — comebacks to K-beauty, straight to your inboxGet it in your inbox›
Persona-Guided LLM Agents for Task-Oriented Dialogue
arXiv:2608.180852026-08-20
AI booking assistants get more likeable when they match a user's personality, but start bending the truth more
Researchers tested what happens when an AI assistant that helps book hotels or restaurants adjusts its responses to match a simulated user's personality, such as being outgoing or anxious. They had two AI agents talk to each other under three conditions: the assistant knowing nothing about the user's personality, guessing it from conversation cues, or being told it outright. Personality-aware responses boosted user satisfaction and how well requests were met, but also made the assistant less truthful.
METAL MEDIA explanatory visual
AI booking assistants get more likeable when they match a user's personality, but start bending the truth more
01The team used three large language models, GPT-4o, Qwen3-Next-80B, and Gemini 2.0 Flash, to simulate a personality-driven 'user' AI talking with a task-completing 'system' AI in hotel and restaurant booking scenarios.
02Three system conditions were compared: Neutral (no personality info), Try (infer personality from dialogue cues), and Oracle (personality given explicitly).
03Both personality-aware conditions improved how well user requirements were satisfied, how much useful information was given, and overall user satisfaction, but reduced truthfulness compared to Neutral.
04When users expressed their assigned trait more strongly, Oracle's satisfaction gains grew accordingly, while Try's gains stayed roughly the same regardless of how clearly the trait showed up, making Try more reliable overall.
05Introverted and closed-off personalities were much harder for the AI to convincingly express, and Qwen3 and Gemini sometimes mirrored a hostile, disagreeable user's tone instead of handling it constructively, unlike GPT-4o.
An explanatory diagram made by METAL MEDIA, not a figure supplied by the paper's authors.
What they did
The team used three large language models, GPT-4o, Qwen3-Next-80B, and Gemini 2.0 Flash, to simulate a personality-driven 'user' AI talking with a task-completing 'system' AI in hotel and restaurant booking scenarios.
Three system conditions were compared: Neutral (no personality info), Try (infer personality from dialogue cues), and Oracle (personality given explicitly).
Both personality-aware conditions improved how well user requirements were satisfied, how much useful information was given, and overall user satisfaction, but reduced truthfulness compared to Neutral.
When users expressed their assigned trait more strongly, Oracle's satisfaction gains grew accordingly, while Try's gains stayed roughly the same regardless of how clearly the trait showed up, making Try more reliable overall.
Introverted and closed-off personalities were much harder for the AI to convincingly express, and Qwen3 and Gemini sometimes mirrored a hostile, disagreeable user's tone instead of handling it constructively, unlike GPT-4o.
Table 1: Trait-level task and personality realization outcomes averaged across domains, models, and the three system personality-access conditions. Dialogue Completion and Task Success summarize schema-guided task performance. User Trait Score measures dialogue-level realization of the intended user trait in generated task-oriented dialogues, while BFI Score measures prompt-level semantic alignment with the intended Big Five direction using a separate questionnaire probe.
Trait
Task Outcomes
Personality Realization
Dialogue Completion (%)
Task Success (%)
User Trait Score
BFI Score
Extraversion
79.67
82.12
4.996
5.000
Agreeableness
95.89
87.44
4.900
4.963
Conscientiousness
93.89
83.34
4.313
5.000
Neuroticism
77.22
75.66
4.971
4.967
Openness
85.33
81.12
4.920
4.967
introversive
96.34
91.34
0.653
4.625
disagreeable
78.78
76.22
4.773
5.000
unconscientious
96.34
95.22
3.169
4.778
stable
97.89
88.22
3.473
4.917
closed
98.11
92.00
0.027
4.667
Average
89.94
85.27
3.620
4.888
Table 2: Trait-level system-quality scores and pairwise user-satisfaction preferences averaged across domains and LLMs. N, T, and O denote Neutral, Try, and Oracle, respectively. CS, IR, and TR denote Constraint Satisfaction, Inform Rate, and Truthfulness, each reported on a 1–5 scale, where higher is better. Pairwise satisfaction cells report win-rate percentages for the first condition over the second; for example, O/N = 62/38 indicates that Oracle is preferred in 62% of comparisons and Neutral in 38%. Values may sum to less than 100 due to ties.
Trait
CS
IR
TR
Pairwise Satisfaction
N
T
O
N
T
O
N
T
O
O/N
T/N
T/O
Extraversion
4.313
4.390
4.307
4.573
4.650
4.543
4.647
4.247
3.887
76.00/24.00
64.33/35.00
40.00/58.66
Agreeableness
4.750
4.897
4.840
4.853
4.960
4.937
4.927
4.893
4.830
71.66/27.34
66.34/31.00
41.33/57.00
Conscientiousness
4.653
4.783
4.790
4.780
4.910
4.917
4.937
4.850
4.883
65.34/34.66
59.00/39.67
46.34/53.33
Neuroticism
4.113
4.263
3.920
4.673
4.757
4.677
4.570
4.257
4.170
63.00/36.67
69.34/30.66
62.66/37.00
Openness
4.213
4.353
4.447
4.763
4.803
4.783
4.547
4.223
3.733
85.00/15.00
60.33/39.34
29.66/70.34
introversive
4.753
4.917
4.880
4.857
4.953
4.963
4.930
4.867
4.890
39.33/59.34
71.00/26.00
73.67/25.33
disagreeable
3.797
4.093
4.063
4.283
4.550
4.553
4.653
4.547
4.657
28.33/71.67
47.00/53.00
80.66/19.34
unconscientious
4.720
4.917
4.883
4.800
4.950
4.927
4.880
4.883
4.853
59.00/40.67
66.67/32.66
65.34/33.66
stable
4.827
4.960
4.920
4.903
4.983
4.963
4.900
4.910
4.873
58.34/40.34
67.00/30.66
53.33/45.00
closed
4.877
4.883
4.933
4.927
4.970
4.983
4.910
4.943
4.977
50.33/49.00
63.00/34.00
68.00/30.34
Average
4.502
4.646
4.598
4.741
4.849
4.825
4.790
4.662
4.575
59.63/39.87
63.40/35.20
56.10/43.00
Table 3: Spearman correlations between Avg. User Trait Score and pairwise satisfaction deltas at two levels of analysis. Trait-level uses one point per trait (n=10), averaged across models and domains. Cell-level uses model × domain × trait cells (n=60, Tables 8–9). p∗<.05; p∗∗∗<.001.
Trait-level (n=10)
Cell-level (n=60)
Delta
𝝆
p
𝝆
p
ΔO-N (Oracle − Neutral)
+.685
.029∗
+.421
<.001∗∗∗
ΔT-N (Try − Neutral)
−.212
.556
−.157
.232
Table 4: Task success rates (%) across models, domains, and personality-access conditions. In the Hotel domain, Neutral achieves the highest cross-model average task success, whereas in the Restaurant domain, Oracle performs best, followed by Try.
Domain / Model
Neutral
Try
Oracle
Hotel
Gemini 2.0 Flash
84.40
78.00
77.40
GPT-4o
94.60
93.40
92.80
Qwen3-Next-80B
95.60
96.20
96.40
Cross-model Average
91.53
89.20
88.87
Restaurant
Gemini 2.0 Flash
70.20
73.60
76.00
GPT-4o
80.00
85.20
88.40
Qwen3-Next-80B
84.00
84.20
84.00
Cross-model Average
78.07
81.00
82.80
Table 5: Dialogue-completion rates (%) for GPT-4o, Qwen3, and Gemini across personality traits in the Hotel and Restaurant domains. N, T, and O denote Neutral, Try, and Oracle, respectively. The average row reports the mean dialogue-completion rate across the ten personality traits for each model and domain.
Model
Trait
Hotel (%)
Restaurant (%)
N
T
O
N
T
O
GPT-4o
Extraversion
96.0
90.0
94.0
82.0
98.0
94.0
Agreeableness
100.0
100.0
98.0
82.0
96.0
98.0
Conscientiousness
96.0
94.0
92.0
86.0
96.0
92.0
Neuroticism
84.0
82.0
70.0
80.0
80.0
72.0
Openness
90.0
80.0
90.0
78.0
88.0
90.0
Introversive
96.0
96.0
96.0
88.0
98.0
98.0
disagreeable
84.0
82.0
84.0
56.0
84.0
78.0
unconscientious
100.0
100.0
100.0
74.0
88.0
98.0
stable
100.0
100.0
100.0
94.0
98.0
94.0
closed
98.0
100.0
98.0
98.0
90.0
96.0
Average
94.4
92.4
92.2
81.8
91.6
91.0
Qwen3
Extraversion
70.0
66.0
54.0
80.0
84.0
82.0
Agreeableness
92.0
90.0
92.0
100.0
100.0
100.0
Conscientiousness
96.0
98.0
98.0
100.0
100.0
100.0
Neuroticism
80.0
84.0
76.0
82.0
88.0
84.0
Openness
88.0
84.0
88.0
90.0
80.0
84.0
Introversive
86.0
94.0
98.0
100.0
100.0
100.0
disagreeable
68.0
74.0
74.0
76.0
78.0
86.0
unconscientious
98.0
100.0
100.0
100.0
100.0
100.0
stable
96.0
100.0
96.0
100.0
100.0
100.0
closed
92.0
98.0
100.0
100.0
100.0
100.0
Average
86.6
88.8
87.6
92.8
93.0
93.6
Gemini
Extraversion
50.0
60.0
52.0
96.0
94.0
92.0
Agreeableness
96.0
96.0
92.0
100.0
98.0
96.0
Conscientiousness
70.0
94.0
90.0
96.0
96.0
96.0
Neuroticism
80.0
72.0
66.0
68.0
66.0
76.0
Openness
80.0
84.0
74.0
94.0
82.0
92.0
Introversive
98.0
100.0
88.0
100.0
100.0
98.0
disagreeable
66.0
70.0
76.0
90.0
92.0
100.0
Table 6: Restaurant-domain system-evaluation results for GPT-4o, Qwen3, and Gemini across personality traits under Neutral (N), Try (T), and Oracle (O) conditions. The table reports Constraint Satisfaction, Truthfulness, and Inform Rate; the average row reports the mean across the ten personality traits for each model.
Model
Trait
Constraint Satisfaction
Truthfulness
Inform Rate
N
T
O
N
T
O
N
T
O
GPT-4o
Extraversion
4.22
4.88
4.94
4.90
4.66
4.78
4.38
5.00
4.98
Agreeableness
4.32
4.84
4.90
4.80
4.58
4.64
4.38
4.84
5.00
Conscientiousness
4.60
4.94
4.76
4.86
4.82
4.82
4.76
4.98
4.84
Neuroticism
4.30
4.40
4.04
4.12
4.02
3.84
4.62
4.68
4.48
Openness
4.26
4.54
4.60
4.64
4.70
4.66
4.44
4.58
4.66
Introversive
4.48
4.90
4.84
4.98
4.76
4.90
4.62
4.92
4.98
disagreeable
3.16
4.12
4.40
4.50
4.28
4.30
3.32
4.56
4.58
unconscientious
4.04
4.76
4.86
4.78
4.92
4.70
4.06
4.76
4.96
stable
4.84
5.00
4.88
4.76
4.76
4.66
4.80
5.00
4.94
closed
4.80
4.68
4.72
5.00
5.00
4.94
4.92
4.88
4.90
Average
4.30
4.71
4.69
4.73
4.65
4.62
4.43
4.82
4.83
Qwen3
Extraversion
4.68
4.70
4.62
4.12
3.72
3.66
4.76
4.78
4.70
Agreeableness
4.96
5.00
5.00
5.00
5.00
4.92
4.96
5.00
5.00
Conscientiousness
5.00
5.00
4.96
4.96
5.00
4.94
5.00
5.00
4.96
Neuroticism
3.92
4.22
3.92
4.74
4.52
4.86
4.68
4.78
4.86
Openness
4.10
4.44
4.52
4.70
4.26
3.40
4.78
4.82
4.76
Introversive
5.00
5.00
5.00
4.94
4.94
4.96
5.00
5.00
5.00
disagreeable
2.78
3.18
3.14
4.38
4.72
4.74
4.16
4.38
4.36
unconscientious
5.00
5.00
5.00
5.00
4.98
5.00
5.00
5.00
5.00
stable
4.94
5.00
4.98
5.00
5.00
5.00
5.00
5.00
5.00
closed
5.00
5.00
5.00
5.00
5.00
5.00
5.00
5.00
5.00
Average
4.54
4.65
4.61
4.78
4.71
4.65
4.83
4.88
4.86
Gemini
Extraversion
4.92
4.94
4.70
4.94
4.68
4.20
4.96
4.96
4.76
Agreeableness
5.00
4.92
4.88
4.96
4.88
4.98
5.00
4.92
4.90
Conscientiousness
4.86
4.92
4.96
5.00
4.64
4.86
4.96
4.96
5.00
Neuroticism
4.08
4.14
4.20
4.92
4.70
4.42
4.62
4.78
4.82
Openness
4.76
4.54
4.72
4.84
4.48
4.48
4.90
4.88
4.86
Introversive
4.98
5.00
4.96
5.00
5.00
4.92
5.00
5.00
5.00
disagreeable
4.64
4.84
4.80
4.96
4.80
4.98
4.66
4.90
5.00
Table 7: Hotel-domain system-evaluation results for GPT-4o, Qwen3, and Gemini across personality traits under Neutral (N), Try (T), and Oracle (O) conditions. The table reports Constraint Satisfaction, Truthfulness, and Inform Rate; the average row reports the mean across the ten personality traits for each model.
Model
Trait
Constraint Satisfaction
Truthfulness
Inform Rate
N
T
O
N
T
O
N
T
O
GPT-4o
Extraversion
4.92
4.76
4.74
4.74
4.36
4.14
4.96
4.82
4.80
Agreeableness
5.00
5.00
5.00
4.96
5.00
4.84
5.00
5.00
5.00
Conscientiousness
4.82
4.68
4.82
4.94
4.92
4.86
4.94
4.84
4.96
Neuroticism
4.42
4.58
4.10
4.62
4.24
4.24
4.84
4.80
4.80
Openness
4.30
4.24
4.58
4.48
3.74
3.36
4.74
4.76
4.88
Introversive
4.82
4.88
4.92
5.00
4.82
4.82
4.90
4.92
5.00
disagreeable
4.68
4.60
4.40
4.82
4.90
4.82
4.86
4.74
4.72
unconscientious
5.00
5.00
4.96
5.00
5.00
4.90
5.00
5.00
5.00
stable
4.96
5.00
5.00
4.94
5.00
4.94
5.00
5.00
5.00
closed
5.00
4.96
4.92
5.00
5.00
5.00
5.00
5.00
5.00
Average
4.79
4.77
4.74
4.85
4.70
4.59
4.92
4.89
4.92
Qwen3
Extraversion
3.76
3.58
3.42
4.18
3.66
2.56
4.68
4.58
4.24
Agreeableness
4.48
4.70
4.52
4.84
4.98
4.82
4.94
5.00
4.88
Conscientiousness
4.66
4.46
4.56
4.94
4.82
4.96
4.96
4.90
4.94
Neuroticism
3.40
3.68
3.22
4.40
3.64
3.26
4.54
4.62
4.42
Openness
3.30
3.70
4.16
3.76
3.82
2.70
4.90
4.86
4.90
Introversive
4.32
4.72
4.88
4.72
4.72
4.88
4.72
4.88
4.96
disagreeable
3.66
3.80
3.52
4.40
3.98
4.28
4.28
4.32
4.34
unconscientious
4.64
4.78
4.84
4.78
4.82
4.82
4.96
4.96
4.94
stable
4.56
4.76
4.84
4.82
4.82
4.84
4.88
4.96
4.92
closed
4.62
4.86
5.00
4.62
4.80
4.98
4.78
4.94
5.00
Average
4.14
4.30
4.30
4.55
4.41
4.21
4.76
4.80
4.75
Gemini
Extraversion
3.38
3.48
3.42
5.00
4.40
3.98
3.70
3.76
3.78
Agreeableness
4.74
4.92
4.74
5.00
4.92
4.78
4.84
5.00
4.84
Conscientiousness
3.98
4.70
4.68
4.92
4.90
4.86
4.06
4.78
4.80
Neuroticism
4.56
4.56
4.04
4.62
4.42
4.40
4.74
4.88
4.68
Openness
4.56
4.66
4.10
4.86
4.34
3.80
4.82
4.92
4.64
Introversive
4.92
5.00
4.68
4.94
4.96
4.86
4.90
5.00
4.84
disagreeable
3.86
4.02
4.12
4.86
4.60
4.82
4.42
4.40
4.32
Table 8: Hotel-domain pairwise user-satisfaction win rates with averaged personality-realization metrics for GPT-4o, Qwen3, and Gemini. O, T, and N denote Oracle, Try, and Neutral, respectively. Each comparison cell reports the win rate of the first condition over the second condition; for example, O/N = 36/58 means Oracle is preferred in 36% of paired comparisons and Neutral in 58%. Percentages may sum to less than 100 when ties occur. BFI Score is the intended-direction score from the BFI probe for the target trait. Avg. User Trait Score and Avg. Rank are averaged across Oracle, Try, and Neutral cases. Lower Avg. Rank indicates stronger realization of the intended trait.
Model
Trait
O / N (%)
T / N (%)
T / O (%)
BFI Score
Avg. User Trait Score
Avg. Rank
GPT-4o
introversive
36/58
64/26
70/24
4.500
0.207
3.913
closed
28/70
50/40
72/26
4.100
0.020
3.987
stable
44/56
60/34
62/36
4.750
3.173
2.653
Extraversion
80/20
64/36
34/62
5.000
4.993
1.000
Agreeableness
76/20
56/36
44/52
4.889
4.887
1.000
Conscientiousness
46/54
44/54
42/56
5.000
4.047
1.327
Neuroticism
54/44
60/40
58/40
5.000
4.920
1.007
Openness
84/16
54/46
28/72
5.000
4.793
1.013
disagreeable
44/56
38/62
66/34
5.000
4.580
1.020
unconscientious
44/54
54/44
64/30
4.444
2.940
2.187
Average
53.6/44.8
54.4/41.8
54.0/43.2
4.768
3.456
1.911
Qwen3
introversive
52/48
56/40
44/56
4.500
0.267
3.873
closed
60/38
56/44
44/56
5.000
0.000
4.000
stable
66/34
54/44
34/66
5.000
3.553
1.833
Extraversion
60/40
58/42
54/46
5.000
5.000
1.000
Agreeableness
76/24
68/32
30/70
5.000
4.820
1.000
Conscientiousness
72/28
48/52
32/68
5.000
4.053
1.053
Neuroticism
70/30
74/26
60/40
5.000
4.993
1.000
Openness
88/12
50/50
16/84
5.000
4.987
1.000
disagreeable
22/78
34/66
74/26
5.000
4.780
1.020
unconscientious
60/40
68/30
54/46
5.000
2.193
2.553
Average
62.6/37.2
56.6/42.6
44.2/55.8
4.950
3.465
1.833
Gemini
introversive
30/70
70/30
86/14
4.875
1.487
3.133
closed
54/46
64/36
74/26
4.900
0.060
3.980
stable
48/50
76/24
64/36
5.000
3.693
1.753
Extraversion
68/32
68/32
44/54
5.000
4.993
1.000
Agreeableness
76/24
70/30
44/56
5.000
4.993
1.000
Conscientiousness
64/36
66/34
64/36
5.000
4.840
1.000
Neuroticism
70/30
84/16
78/22
5.000
5.000
1.000
Openness
92/8
92/8
34/66
4.900
4.980
1.007
Table 9: Restaurant-domain pairwise user-satisfaction win rates with averaged personality-realization metrics for GPT-4o, Qwen3, and Gemini. O, T, and N denote Oracle, Try, and Neutral, respectively. Each comparison cell reports the win rate of the first condition over the second condition, e.g., O/N = 62/38 means Oracle is preferred in 62% of paired comparisons and Neutral in 38%. Percentages may sum to less than 100 when ties occur. BFI Score is the intended-direction score from the BFI probe for the target trait. Avg. User Trait Score and Avg. Rank are averaged across Oracle, Try, and Neutral cases. Lower Avg. Rank indicates stronger realization of the intended trait.
Model
Trait
O / N (%)
T / N (%)
T / O (%)
BFI Score
Avg. User Trait Score
Avg. Rank
GPT-4o
introversive
62/38
82/18
60/40
4.500
1.420
3.173
closed
54/46
58/36
60/34
4.100
0.393
3.760
stable
66/34
76/24
62/36
4.750
3.260
2.213
Extraversion
88/12
66/32
28/70
5.000
4.973
1.000
Agreeableness
76/24
78/18
38/56
4.889
4.593
1.020
Conscientiousness
70/30
70/30
54/46
5.000
3.507
1.807
Neuroticism
50/50
58/42
62/38
5.000
4.673
1.020
Openness
80/20
60/38
40/60
5.000
4.827
1.020
disagreeable
70/30
64/36
64/36
5.000
4.527
1.000
unconscientious
70/30
78/22
68/32
4.444
1.260
3.413
Average
68.6/31.4
69.0/29.6
53.6/44.8
4.768
3.343
1.943
Qwen3
introversive
4/96
70/26
98/2
4.500
0.240
3.867
closed
42/58
72/26
84/16
5.000
0.000
4.000
stable
76/22
70/26
38/60
5.000
3.580
1.747
Extraversion
72/28
62/36
48/52
5.000
4.987
1.000
Agreeableness
88/10
68/28
18/82
5.000
4.787
1.013
Conscientiousness
74/26
76/18
38/62
5.000
4.073
1.020
Neuroticism
50/50
68/32
70/30
5.000
4.980
1.000
Openness
82/18
52/48
20/80
5.000
4.913
1.040
disagreeable
20/80
58/42
86/14
5.000
4.960
1.000
unconscientious
60/40
70/30
72/28
5.000
2.027
2.660
Average
56.8/42.8
66.6/31.2
57.2/42.6
4.950
3.455
1.835
Gemini
introversive
52/46
84/16
84/16
4.875
1.895
2.969
closed
64/36
78/22
74/24
4.900
0.140
3.940
stable
50/46
66/32
60/36
5.000
3.660
1.820
Extraversion
88/12
68/32
32/68
5.000
5.000
1.000
Agreeableness
38/62
58/42
74/26
5.000
5.000
1.000
Conscientiousness
66/34
50/50
48/52
5.000
4.600
1.000
Neuroticism
84/16
72/28
48/52
5.000
5.000
1.000
Openness
84/16
54/46
40/60
4.900
4.953
1.087
Table 10: Per-trait Avg. Intended Score and average pairwise satisfaction deltas (pp), averaged across six model–domain combinations. Traits are sorted by Avg. Intended Score in descending order. The disagreeable trait is discussed separately in Appendix C.3.
Trait
Avg. Intended Score
Avg. ΔO−N (%)
Avg. ΔT−N (%)
Extraversion
4.991
+52.0
+29.3
Neuroticism
4.928
+26.3
+38.7
Openness
4.909
+70.0
+21.0
Agreeableness
4.847
+44.3
+35.3
disagreeable†
4.788
−43.3
−6.0
Conscientiousness
4.187
+30.7
+19.3
stable
3.486
+18.0
+36.3
unconscientious
2.737
+18.3
+34.0
introversive
0.919
−20.0
+45.0
closed
0.102
+1.3
+29.0
Why it matters
This shows a concrete tradeoff for anyone building customer-facing chat assistants: tailoring tone to a user's personality can raise satisfaction but risks the assistant fabricating or distorting information. It also demonstrates that this kind of personality-aware behavior can be achieved through prompting alone, without retraining the underlying model.
Terms in this paper
Big Five traits · a psychology framework describing personality across five dimensions: extraversion, agreeableness, conscientiousness, neuroticism, and openness
task-oriented dialogue (TOD) · multi-turn conversation aimed at completing a concrete goal, like booking a table
Constraint Satisfaction · a metric checking whether the final recommendation or booking actually meets the user's stated requirements
Truthfulness · a metric checking whether the system avoids making up unsupported claims and stays grounded in real dialogue and data
Spearman correlation · a rank-based statistic measuring whether two variables tend to rise and fall together
Figures we cannot republish
Figure 1: Overview of our framework for personality-aware task-oriented dialogue. A personality-conditioned user agent interacts with a system agent under three personality-access conditions—Neutral, Try, and Oracle—within an SGD-based API-grounded task setting. Generated dialogues are evaluated for task outcomes, system quality, and personality-related behavior.
Figure 2: Spearman correlation between Avg. User Intended Score and pairwise satisfaction deltas across traits (n=10). Oracle gains increase with trait realization (ρ=+0.685, p=.029), whereas Try gains do not (ρ=−0.103, p=.777). Point size indicates cross-dataset SD; disagreeable reflects a Qwen3/Gemini mirroring artifact (Appendix C.3).
Figure 3: Representative GPT-4o Restaurant-domain score-delta figures. The top panel shows score changes from Neutral to Try, and the bottom panel shows score changes from Neutral to Oracle.
Prior work has shown that large language models (LLMs) can express diverse personality traits in open-ended text generation. However, it remains unclear whether they can do so in a goal-directed dialogue without compromising task completion, and whether adapting to the user's personality improves the interaction quality. We study these questions in task-oriented dialogue (TOD), where a system helps a user accomplish a goal via multi-turn interaction. We build a training-free framework that simulates a TOD interaction between two LLMs: a user agent that exhibits a target personality and a system agent that adapts to the user while completing the task. To isolate the effect of adaptation, we vary how much the system knows about the user's personality across three conditions. In Neutral, the system receives no personality information. In Try, it infers the personality from dialogue cues. In Oracle, it is given the personality explicitly. We evaluate GPT-4o, Qwen3-Next-80B, and Gemini 2.0 Flash on Hotel and Restaurant dialogues from the Schema-Guided Dialogue (SGD) dataset, across the Big Five traits and their opposite poles. We find that the user agent can express personality while the system maintains strong task performance, although some traits are realized far less reliably than others. Adapting to the user's personality improves constraint satisfaction, inform rate, and user satisfaction, but lowers truthfulness, revealing a trade-off between personalization and task-grounding. Oracle's gains grow when the target trait is strongly expressed, whereas Try's gains are largely insensitive to realization strength. Overall, cue-based adaptation in Try best resolves this trade-off and offers a more reliable route to personality-aware TOD without fine-tuning.
Authors · Maryam Shoaeinaeini, Brent Harrison, A. B. Siddique