Skip to main content
QUICK REVIEW

[论文解读] Pragmatically Appropriate Diversity for Dialogue Evaluation

Katherine Stasaski, Marti A. Hearst|arXiv (Cornell University)|Apr 6, 2023
Topic Modeling被引用 5
一句话总结

本文提出了实用得当多样性(PA Diversity)框架,该框架基于前一句话语的言语行为来评估对话回复的多样性。研究显示,自动多样性度量与创意写作者的人工判断均与PA Diversity一致,表明回复多样性应随言语行为类型而变化,而非在所有对话中统一处理。

ABSTRACT

Linguistic pragmatics state that a conversation's underlying speech acts can constrain the type of response which is appropriate at each turn in the conversation. When generating dialogue responses, neural dialogue agents struggle to produce diverse responses. Currently, dialogue diversity is assessed using automatic metrics, but the underlying speech acts do not inform these metrics. To remedy this, we propose the notion of Pragmatically Appropriate Diversity, defined as the extent to which a conversation creates and constrains the creation of multiple diverse responses. Using a human-created multi-response dataset, we find significant support for the hypothesis that speech acts provide a signal for the diversity of the set of next responses. Building on this result, we propose a new human evaluation task where creative writers predict the extent to which conversations inspire the creation of multiple diverse responses. Our studies find that writers' judgments align with the Pragmatically Appropriate Diversity of conversations. Our work suggests that expectations for diversity metric scores should vary depending on the speech act.

研究动机与目标

  • 探究对话中的言语行为类型是否限制了合适回复的多样性。
  • 开发一种新的评估度量方法——实用得当多样性(PA Diversity),以考虑基于言语行为的回复多样性约束。
  • 通过一项新颖的人工评估任务,邀请创意写作者对对话提示中潜在的多样化回复进行评分,以验证PA Diversity。
  • 评估自动多样性度量(如NLI Diversity)是否能检测出基于言语行为类型的多样性差异。
  • 通过展示多样性期望应随言语行为类型而变化,为未来对话模型的评估与生成提供指导。

提出的方法

  • 作者分析了一个包含多条回复的对话数据集(DailyDialog++),该数据集包含人工标注和自动预测的言语行为。
  • 在不同言语行为类型的回复集合上计算自动多样性度量——特别是NLI Diversity和基于Sentence-BERT的多样性。
  • 设计了一项新的人工评估任务,邀请创意写作者在5分制量表上评估对话提示激发多种多样化回复的潜力。
  • 将写作者的判断与自动度量及言语行为类型进行比较,以验证PA Diversity概念。
  • 使用统计分析检验PA Diversity得分在不同言语行为类别间是否存在显著差异。
  • 探索利用PA Diversity指导模型评估与生成的潜力,提出针对低PA-Diversity对话的基于规则的系统。
Figure 2: NLI Diversity (top) and Sent-BERT (bottom) for responses categorized by most-recent speech act utterance (higher values indicate more diverse, ordered by diversity). Mean values are indicated by the white circle and corresponding text label. Box-and-whisker plots show the interquartile ran
Figure 2: NLI Diversity (top) and Sent-BERT (bottom) for responses categorized by most-recent speech act utterance (higher values indicate more diverse, ordered by diversity). Mean values are indicated by the white circle and corresponding text label. Box-and-whisker plots show the interquartile ran

实验结果

研究问题

  • RQ1对话中最近一句的言语行为是否限制了实用得当回复的多样性?
  • RQ2自动多样性度量能否检测出基于言语行为类型的回复多样性差异?
  • RQ3创意写作者的人工判断是否与对话的实用得当多样性一致?
  • RQ4NLI Diversity在区分基于言语行为的多样性差异方面是否优于Sentence-BERT?
  • RQ5PA Diversity能否用于指导对话模型的评估与生成策略?

主要发现

  • 最近一句的言语行为显著限制了合适回复的多样性,其中问题类回复的多样性高于道歉或结束语类。
  • 自动多样性度量(尤其是NLI Diversity)能有效检测出不同言语行为类型间回复多样性的有意义差异。
  • 创意写作者对回复多样性的评分与PA Diversity假设存在显著相关性,支持该假设的有效性。
  • NLI Diversity在区分不同言语行为的多样性水平方面优于Sentence-BERT,证实其对语义差异的敏感性。
  • 本研究发现,并非所有对话在回复潜力上都具有同等多样性,多样性期望应根据言语行为类型进行上下文调整。
  • 部分写作者对所有对话均给出高分,表明其对创作潜力的感知存在个体差异,凸显了上下文感知评估的必要性。
Pragmatically Appropriate Diversity for Dialogue Evaluation

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。