Skip to main content
QUICK REVIEW

[論文レビュー] DICES Dataset: Diversity in Conversational AI Evaluation for Safety

Lora Aroyo, Alex Taylor|arXiv (Cornell University)|Jun 20, 2023
Computational and Text Analysis Methods被引用数 6
ひとこと要約

DICESデータセットは、会話型AIの安全性を評価するための大規模で文化的に多様なベンチマークを導入し、1,340件の会話(人間とボット間)に対して250万件以上の評価を収集。評価者には細分化された人種的・文化的背景が明示されており、1件の会話あたり70〜120件の評価が得られており、再現性が極めて高い。このデータセットにより、主観的な安全性認識、評価者の合意不一致、および多様性が安全性判断に与える影響を詳細に分析可能となり、従来の「ゴールドラベル」仮定に疑問を呈し、包摂的かつバイアスに配慮したモデル評価の基盤を提供する。

ABSTRACT

Machine learning approaches often require training and evaluation datasets with a clear separation between positive and negative examples. This risks simplifying and even obscuring the inherent subjectivity present in many tasks. Preserving such variance in content and diversity in datasets is often expensive and laborious. This is especially troubling when building safety datasets for conversational AI systems, as safety is both socially and culturally situated. To demonstrate this crucial aspect of conversational AI safety, and to facilitate in-depth model performance analyses, we introduce the DICES (Diversity In Conversational AI Evaluation for Safety) dataset that contains fine-grained demographic information about raters, high replication of ratings per item to ensure statistical power for analyses, and encodes rater votes as distributions across different demographics to allow for in-depth explorations of different aggregation strategies. In short, the DICES dataset enables the observation and measurement of variance, ambiguity, and diversity in the context of conversational AI safety. We also illustrate how the dataset offers a basis for establishing metrics to show how raters' ratings can intersects with demographic categories such as racial/ethnic groups, age groups, and genders. The goal of DICES is to be used as a shared resource and benchmark that respects diverse perspectives during safety evaluation of conversational AI systems.

研究の動機と目的

  • 安全性が文化的・社会的に根ざしていることを踏まえ、会話型AIの安全性評価において多様で代表的な視点が欠落している問題に対処すること。
  • 安全性に関する評価のばらつき、曖昧さ、多様性を捉えることで、単純な「真の正解」ラベルの枠組みを越えること。
  • 人種/民族、年齢、性別などのデモグラフィック要因が安全性判断に与える影響を統計的パワーと詳細分析を可能にするベンチマークを提供すること。
  • 単一の基準を強制するのではなく、安全性の多様な見方を反映したファインチューニングおよび評価手法の開発を可能にすること。
  • 高水準の評価者合意不一致を示すことで、従来の「ゴールドラベル」の概念に疑問を呈し、データセットの品質とモデル学習への影響を検証すること。

提案手法

  • 性別、年齢、人種/民族的背景の各グループに配分されたバランスの取れた大規模な評価者プールから、安全性評価を収集することで、デモグラフィックの代表性を確保。
  • DICES-990(990件の会話)およびDICES-350(350件の会話)の各会話に対して、70〜120件の評価が得られ、統計的パワーが高く、ばらつきの推定に適したリサンプリングが可能。
  • 有害性、バイアス、誤情報、政治的発言、ポリシー違反の5つの安全性カテゴリーに分類され、ハートスピーチなどの特定タイプのサブレーティングも含まれる。
  • 評価者の投票をデモグラフィックグループごとの分布としてエンコードし、多段階ベイジアンモデリングを用いて合意不一致や交差的影響を分析可能。
  • 安全性、有害性、有害度の専門家アノテーションを含め、クラウド評価と専門家判断の比較が可能。
  • 時間的・行動的データを活用し、外れ値やノイズを検出するための合意不一致指標と評価者行動の分析を支援。
Figure 2: Screenshot of the raters’ user interface for the Safety Annotation Task: illustrates the annotation category for policy violations . The left panel presents the conversation; raters assess the last conversational turn (highlighted). The right panel presents two policy related sub-questions
Figure 2: Screenshot of the raters’ user interface for the Safety Annotation Task: illustrates the annotation category for policy violations . The left panel presents the conversation; raters assess the last conversational turn (highlighted). The right panel presents two policy related sub-questions

実験結果

リサーチクエスチョン

  • RQ1人種、年齢、性別などの異なるデモグラフィックグループは、会話型AI出力の安全性認識にどのように差を示すか?
  • RQ2評価者の多様性が安全性判断における合意不一致をどの程度引き起こすか。その統計的モデリングは可能か?
  • RQ3交差的アイデンティティ(例:若年層のブラック・ウィメン)は、単一軸のグループ分けと比較して、安全性認識にどのような影響を及えるか?
  • RQ4評価者間の合意不一致が顕著な状況下でも、専門家アノテーションとクラウド評価を意味的に比較可能か?
  • RQ5高レプリカション率が、ばらつきの推定およびモデル評価の耐性向上に与える影響は何か?

主な発見

  • DICESデータセットには、1,340件の会話に対して、約300人の評価者から250万件以上の評価が収集されており、1件あたり70〜120件の評価が得られており、通常の3〜5件の基準を大きく上回っている。
  • 評価者間の合意不一致が顕著で、安全性評価に顕著な主観性があることが示され、安全性のための単一の「ゴールドラベル」の実現可能性に疑問を呈する。
  • すべてのデモグラフィックグループが安全性認識に同じ程度の影響を及ぼすわけではない。一部のサブグループでは、より強く一貫性のある評価パターンが見られる。
  • 交差的グループ(例:若年層のブラック・ウィメン)は、明確な安全性認識のパターンを示しており、単一軸のデモグラフィック分析では見逃されがちな重要な洞察を明らかにする。
  • 高レプリカション率のおかげで、リサンプリングが安定し、ばらつきの推定がより正確に可能となり、安全性評価研究における統計的推論の信頼性が向上する。
  • 専門家評価とクラウド評価には乖離が認められ、単一の専門家コンSENSUSに依存するのではなく、多様な視点を統合する手法の必要性が浮き彫りになる。
Figure 3: Breakdown of topics and degree of harm for DICES-350. Percentages of conversations per topic (left) and number of conversations per degree of harm (right).
Figure 3: Breakdown of topics and degree of harm for DICES-350. Percentages of conversations per topic (left) and number of conversations per degree of harm (right).

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。