Generating Diverse Personas for User Simulators to Test Interview Dialogue Systems
要测试访谈式对话系统需要大量不同性格的虚拟用户,这项研究用大语言模型自动生成这些虚拟用户人设
询问旅行计划或甜点喜好等信息的访谈类对话系统,用真人测试成本很高。这篇论文只需几个人工写的示例人设,就能让大语言模型自动批量生成风格各异的虚拟用户人设,再用这些人设驱动模拟用户进行对话测试。实验显示,这样生成的模拟对话比只用固定人设时更加多样化。
METAL MEDIA 解读图
要测试访谈式对话系统需要大量不同性格的虚拟用户,这项研究用大语言模型自动生成这些虚拟用户人设
- 01方法在两个日语访谈对话系统上做了测试,一个询问旅行相关信息,一个询问甜点偏好
- 02从10个人工编写的种子人设出发,用少样本上下文学习的方式让GPT-4o为每种条件生成100个新人设
- 03生成人设时还指定了两种与沟通风格相关的性格特质:拟人化程度(把系统当物品还是当人对待)和表达详略程度(说话是绕弯子还是直接)
- 04用生成的人设驱动基于GPT-4o的模拟用户,与基于GPT-4o-mini的访谈系统对话,再用发言长度差异、词汇型符比等指标衡量多样性
- 05仅靠大语言模型生成新人设就已经提升了内容多样性(实词型符比在旅行领域从.106升到.122,在甜点领域从.109升到.133),而加入详略程度这一特质后,发言长度的标准差也明显提升(旅行领域从7.0升到18.2,甜点领域从8.0升到17.7),说明文体多样性也增加了
他们做了什么
- 方法在两个日语访谈对话系统上做了测试,一个询问旅行相关信息,一个询问甜点偏好
- 从10个人工编写的种子人设出发,用少样本上下文学习的方式让GPT-4o为每种条件生成100个新人设
- 生成人设时还指定了两种与沟通风格相关的性格特质:拟人化程度(把系统当物品还是当人对待)和表达详略程度(说话是绕弯子还是直接)
- 用生成的人设驱动基于GPT-4o的模拟用户,与基于GPT-4o-mini的访谈系统对话,再用发言长度差异、词汇型符比等指标衡量多样性
- 仅靠大语言模型生成新人设就已经提升了内容多样性(实词型符比在旅行领域从.106升到.122,在甜点领域从.109升到.133),而加入详略程度这一特质后,发言长度的标准差也明显提升(旅行领域从7.0升到18.2,甜点领域从8.0升到17.7),说明文体多样性也增加了
| Condition | Personality | #Dialogues | Ave. utterance length (S.D.) | Ave. utterance | length (S.D.) | Total words | Total | words | Unique words | Unique | words | Unique bigrams | Unique | bigrams | TTR |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Ave. utterance | |||||||||||||||
| length (S.D.) | |||||||||||||||
| Total | |||||||||||||||
| words | |||||||||||||||
| Unique | |||||||||||||||
| words | |||||||||||||||
| Unique | |||||||||||||||
| bigrams | |||||||||||||||
| BL | 100 | 28.2 | (7.7) | 42,260 | 15,747 | 31,620 | .373 | ||||||||
| noPT | 100 | 28.5 | (7.0) | 42,784 | 16,362 | 32,730 | .382 | ||||||||
| APM | All | 100 | 28.4 | (7.2) | 42,563 | 16,089 | 32,334 | .378 | |||||||
| High | 50 | 30.1 | (7.6) | 22,547 | 8,406 | 17,098 | .373 | ||||||||
| Low | 50 | 26.7 | (6.3) | 20,016 | 7,683 | 15,236 | .384 | ||||||||
| EL | All | 100 | 36.0 | (18.2) | 53,951 | 18,556 | 39,010 | .344 | |||||||
| High | 50 | 50.7 | (13.9) | 38,001 | 12,175 | 26,883 | .320 | ||||||||
| Low | 50 | 21.3 | (5.7) | 15,950 | 6,381 | 12,127 | .400 | ||||||||
| APM+EL | All | 100 | 31.4 | (13.4) | 47,111 | 16,856 | 34,839 | .358 | |||||||
| High+High | 25 | 46.2 | (12.3) | 17,308 | 5,624 | 12,309 | .325 | ||||||||
| High+Low | 25 | 23.6 | (4.8) | 8,836 | 3,451 | 6,787 | .391 | ||||||||
| Low+High | 25 | 35.2 | (9.9) | 13,184 | 4,651 | 9,796 | .353 | ||||||||
| Low+Low | 25 | 20.8 | (5.9) | 7,783 | 3,130 | 5,947 | .402 |
| Condition | Personality | #Dialogues | Ave. utterance length (S.D.) | Ave. utterance | length (S.D.) | Total words | Total | words | Unique words | Unique | words | Unique bigrams | Unique | bigrams | TTR |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Ave. utterance | |||||||||||||||
| length (S.D.) | |||||||||||||||
| Total | |||||||||||||||
| words | |||||||||||||||
| Unique | |||||||||||||||
| words | |||||||||||||||
| Unique | |||||||||||||||
| bigrams | |||||||||||||||
| BL | 100 | 25.9 | (7.8) | 25,152 | 11,138 | 20,534 | .443 | ||||||||
| noPT | 100 | 25.6 | (8.0) | 25,509 | 11,105 | 20,634 | .435 | ||||||||
| APM | All | 100 | 25.6 | (8.2) | 25,680 | 11,124 | 20,862 | .433 | |||||||
| High | 50 | 26.5 | (8.2) | 13,269 | 5,723 | 10,824 | .431 | ||||||||
| Low | 50 | 24.6 | (8.0) | 12,411 | 5,401 | 10,038 | .435 | ||||||||
| EL | All | 100 | 33.9 | (17.7) | 34,020 | 13,214 | 26,117 | .388 | |||||||
| High | 50 | 47.5 | (14.8) | 23,755 | 8,578 | 17,808 | .361 | ||||||||
| Low | 50 | 20.3 | (6.1) | 10,265 | 4,636 | 8,309 | .452 | ||||||||
| APM+EL | All | 100 | 28.1 | (12.2) | 28,067 | 11,549 | 22,182 | .411 | |||||||
| High+High | 25 | 39.4 | (13.1) | 9,857 | 3,743 | 7,643 | .380 | ||||||||
| High+Low | 25 | 22.2 | (6.7) | 5,560 | 2,458 | 4,533 | .442 | ||||||||
| Low+High | 25 | 30.6 | (10.5) | 7,642 | 3,087 | 6,015 | .404 | ||||||||
| Low+Low | 25 | 20.0 | (5.9) | 5,008 | 2,261 | 3,991 | .451 |
为什么重要
开发者不必招募真人测试者,就能用这种方法对访谈对话系统进行覆盖多种用户行为的压力测试,降低开发成本和人力投入。人设越多样,发现系统未曾预料到的问题的概率就越高。
本文术语
- 人设(persona) · 赋予虚拟用户的性格、偏好和说话风格等信息
- 访谈对话系统 · 通过提问从用户那里收集信息的对话式人工智能系统
- 用户模拟器 · 代替真人与对话系统进行交互的虚拟对话对象
- 型符比(TTR) · 不同词数量占总词数的比例,用来衡量词汇多样性的指标
- 少样本上下文学习 · 在提示中给模型看几个示例,让它据此生成风格相似的新内容
论文原文摘要(英文)
This paper addresses the issue of the significant labor required to test interview dialogue systems. While interview dialogue systems are expected to be useful in various scenarios, like other dialogue systems, testing them with human users requires significant effort and cost. Therefore, testing with user simulators can be beneficial. Since most conventional user simulators have been primarily designed for training task-oriented dialogue systems, little attention has been paid to the personas of the simulated users. During development, testing interview dialogue systems requires simulating a wide range of user behaviors, but manually creating a large number of personas is labor-intensive. We propose a method that automatically generates personas for user simulators using a large language model. Furthermore, by assigning personality traits related to communication styles when generating personas, we aim to increase the diversity of communication styles in the user simulator. Experimental results show that the proposed method enables the user simulator to generate utterances with greater variation.
在 arXiv 阅读最新论文
- SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?让AI编程助手去修复真实科学软件,连最强的那个也有一半以上任务没做对
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving把稀疏注意力从论文原型变成能真正上线服务的加速方案
- PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents让客服AI坐席不只是拦住一个危险动作,而是把整个流程走对
- EXIMO: VLM Guided Exploration of VLA Policies不用人工遥控演示,让会说话的AI来教机械臂做新家务
- EnvHarness: Awakening Static Worlds for Agent Learning不重新搭建训练环境,而是给现有环境套一层可插拔组件,针对每个智能体的具体弱点重新塑形
- Bounded Sovereignty and the Control Tax: Pricing AI Oversight When the Deployer Does Not Own the Model租用AI而非拥有AI的机构,安全监管能力只剩一半
- Beyond Imitation: Filtering On-Policy Distillation by Reasoning ProgressAI模仿老师模型学习时,会误伤本来推理正确的步骤,新方法专门过滤掉这种误伤
- PersonalBench: Measuring the Authorship Gap in LLM Personalization让AI模仿某人的文风,结果发现它始终摆脱不了自己的腔调
METAL MEDIA 最新报道
图片来源: Mikio Nakano et al., arXiv:2608.19549, arxiv-nonexclusive