[论文解读] Adding Chit-Chats to Enhance Task-Oriented Dialogues
本文提出 ACCENTOR,一种通过整合人工参与的对话收集机制所获得的闲聊回复,来增强任务导向对话的框架,使虚拟助手能够自然地在任务执行与社交对话之间切换。在 Schema-Guided Dialogue 和 MultiWOZ 2.1 上的评估表明,该方法在不牺牲任务性能的前提下,提升了对话的参与度和自然度。
Existing dialogue corpora and models are typically designed under two disjoint motives: while task-oriented systems focus on achieving functional goals (e.g., booking hotels), open-domain chatbots aim at making socially engaging conversations. In this work, we propose to integrate both types of systems by Adding Chit-Chat to ENhance Task-ORiented dialogues (ACCENTOR), with the goal of making virtual assistant conversations more engaging and interactive. Specifically, we propose a Human AI collaborative data collection approach for generating diverse chit-chat responses to augment task-oriented dialogues with minimal annotation effort. We then present our new chit-chat-based annotations to 23.8K dialogues from two popular task-oriented datasets (Schema-Guided Dialogue and MultiWOZ 2.1) and demonstrate their advantage over the originals via human evaluation. Lastly, we propose three new models for adding chit-chat to task-oriented dialogues, explicitly trained to predict user goals and to generate contextually relevant chit-chat responses. Automatic and human evaluations show that, compared with the state-of-the-art task-oriented baseline, our models can code-switch between task and chit-chat to be more engaging, interesting, knowledgeable, and humanlike, while maintaining competitive task performance.
研究动机与目标
- 通过将闲聊整合到功能性对话中,弥合任务导向对话系统与开放域聊天机器人之间的差距。
- 通过一种人工与 AI 协同的数据收集方法,减少收集多样化闲聊回复的标注工作量。
- 在保持任务成功率的前提下,提升虚拟助手对话的自然度、参与度和自然感。
- 开发能够根据用户目标和上下文动态切换任务导向与闲聊回复的模型。
提出的方法
- 采用人工与 AI 协同的数据收集流程,以极低的人工标注成本生成多样化闲聊回复,用于任务导向对话的各个回合。
- 在 Schema-Guided Dialogue 和 MultiWOZ 2.1 的 23.8K 个对话中对收集到的闲聊回复进行了标注,丰富了原始数据集。
- 训练了三种新型神经模型,以联合预测用户目标并生成上下文相关的闲聊回复。
- 模型采用端到端方式训练,以根据对话上下文动态切换任务导向与闲聊模式。
- 通过人类评估验证生成的闲聊回复的质量与相关性。
- 通过自动评估与人类评估,将所提出的模型与当前最先进的任务导向基线模型进行对比。
实验结果
研究问题
- RQ1闲聊回复是否能在不降低任务性能的前提下,提升任务导向对话的参与度与自然感?
- RQ2如何在大规模下高效地收集闲聊内容,同时保持多样性与相关性?
- RQ3模型在多大程度上能够基于上下文与用户目标学习在任务模式与闲聊模式之间进行代码切换?
- RQ4与基线模型相比,人工标注的闲聊回复在质量与自然度方面表现如何?
- RQ5在联合目标预测与闲聊生成任务上端到端训练的模型,是否优于独立的任务与回复生成系统?
主要发现
- 人类评估确认,加入闲聊回复的对话在参与度、趣味性和自然感方面,显著优于原始的任务导向对话。
- 所提出的模型在保持高成功率的同时,实现了与基线模型相当的任务性能,且在 Schema-Guided Dialogue 和 MultiWOZ 2.1 上均表现良好。
- 人工与 AI 协同的数据收集方法在减少标注工作量的同时,生成了多样化且上下文相关的闲聊回复。
- 采用联合目标预测与闲聊生成训练的模型,在任务与社交模式之间表现出更优的代码切换行为。
- 自动评估显示,与基线系统相比,该模型生成的闲聊回复在上下文相关性与多样性方面表现更优。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。