Skip to main content
QUICK REVIEW

[論文レビュー] Pragmatically Appropriate Diversity for Dialogue Evaluation

Katherine Stasaski, Marti A. Hearst|arXiv (Cornell University)|Apr 6, 2023
Topic Modeling被引用数 5
ひとこと要約

本稿では、前回の発話の言語行為に基づいて応答の多様性を評価するフレームワークであるPragmatically Appropriate Diversity(PA Diversity)を紹介する。自動多様性メトリクスとクリエイティブライターの人的 judgmentsがPA Diversityと一致しており、応答の多様性は会話の種類に応じて変化すべきであり、会話全体に一様に扱うべきでないことが示された。

ABSTRACT

Linguistic pragmatics state that a conversation's underlying speech acts can constrain the type of response which is appropriate at each turn in the conversation. When generating dialogue responses, neural dialogue agents struggle to produce diverse responses. Currently, dialogue diversity is assessed using automatic metrics, but the underlying speech acts do not inform these metrics. To remedy this, we propose the notion of Pragmatically Appropriate Diversity, defined as the extent to which a conversation creates and constrains the creation of multiple diverse responses. Using a human-created multi-response dataset, we find significant support for the hypothesis that speech acts provide a signal for the diversity of the set of next responses. Building on this result, we propose a new human evaluation task where creative writers predict the extent to which conversations inspire the creation of multiple diverse responses. Our studies find that writers' judgments align with the Pragmatically Appropriate Diversity of conversations. Our work suggests that expectations for diversity metric scores should vary depending on the speech act.

研究の動機と目的

  • 会話における言語行為の種別が、適切な応答の多様性に制約を及えるかどうかを調査すること。
  • 言語行為に基づく制約を反映した新しい評価指標、Pragmatically Appropriate Diversity(PA Diversity)を開発すること。
  • クリエイティブライターが会話のプロンプトが提示する多様な応答の可能性を評価する新しい人的評価タスクを用いて、PA Diversityを検証すること。
  • NLI Diversityなどの自動多様性メトリクスが、言語行為の種別に応じた多様性の違いを検出できるかどうかを評価すること。
  • 言語行為に応じた多様性の期待値の違いを示すことで、将来の会話モデルの評価と生成を支援すること。

提案手法

  • 著者らは、人間ラベルと自動予測の両方の言語行為ラベルが付与されたマルチ・リスポンス会話データセット(DailyDialog++)を分析した。
  • 異なる言語行為の種別ごとに応答集合に対して自動多様性メトリクス(特にNLI DiversityとSentence-BERTベースの多様性)を計算した。
  • クリエイティブライターが会話プロンプトが複数の多様な応答を生み出す可能性を5段階スケールで評価する新しい人的評価タスクを設計した。
  • ライターの判断と自動メトリクス、言語行為の種別を比較し、PA Diversityの概念を検証した。
  • 統計的分析を用いて、PA Diversityスコアが言語行為のカテゴリ間で有意に異なるかどうかを検証した。
  • PA Diversityをモデルの評価と生成戦略の指針とする可能性を検討し、低PA-Diversity会話に対してルールベースのシステムを提案した。
Figure 2: NLI Diversity (top) and Sent-BERT (bottom) for responses categorized by most-recent speech act utterance (higher values indicate more diverse, ordered by diversity). Mean values are indicated by the white circle and corresponding text label. Box-and-whisker plots show the interquartile ran
Figure 2: NLI Diversity (top) and Sent-BERT (bottom) for responses categorized by most-recent speech act utterance (higher values indicate more diverse, ordered by diversity). Mean values are indicated by the white circle and corresponding text label. Box-and-whisker plots show the interquartile ran

実験結果

リサーチクエスチョン

  • RQ1会話における最新の言語行為が、意味的に適切な応答の多様性を制約するか?
  • RQ2自動多様性メトリクスは、言語行為の種別に応じた応答多様性の違いを検出できるか?
  • RQ3クリエイティブライターの人的 judgmentsは、会話のPragmatically Appropriate Diversityと一致するか?
  • RQ4NLI DiversityはSentence-BERTに比べて、言語行為に基づく多様性の違いをより効果的に検出できるか?
  • RQ5PA Diversityは、会話モデルの評価と生成戦略の指針として利用可能か?

主な発見

  • 最新の言語行為が、適切な応答の多様性に顕著な制約を及えることが判明した。質問は謝罪や終了発話よりも高い多様性を許容する。
  • 自動多様性メトリクス、特にNLI Diversityは、言語行為の種別に応じた応答多様性の有意な差を検出できた。
  • クリエイティブライターの応答多様性評価はPA Diversity仮説と有意に相関しており、その妥当性を支持した。
  • NLI Diversityは、意味的変動に敏感であるため、言語行為の種別に応じた多様性の違いをSentence-BERTよりも効果的に区別できた。
  • 本研究では、すべての会話が応答の可能性において均一に多様ではないことが明らかになり、多様性の期待値は会話の文脈(言語行為)に応じて調整すべきであることが示された。
  • 一部のライターは一貫してすべての会話を高く評価しており、創造的潜在力の認識に個人差があることが示された。これは、文脈に配慮した評価の必要性を強調している。
Pragmatically Appropriate Diversity for Dialogue Evaluation

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。