Skip to main content
QUICK REVIEW

[논문 리뷰] DICES Dataset: Diversity in Conversational AI Evaluation for Safety

Lora Aroyo, Alex Taylor|arXiv (Cornell University)|2023. 06. 20.
Computational and Text Analysis Methods인용 수 6
한 줄 요약

DICES 데이터셋은 대규모이고 민족·인구통계학적으로 다양한 기준을 바탕으로 대화형 AI의 안전성 평가를 위한 벤치마크를 제공하며, 1,340개의 인간-봇 대화에서 250만 건 이상의 평가를 기록하고, 평가자들의 세밀한 인구통계학적 정보와 높은 복제 수(1회 대화당 70~120건)를 확보하였다. 이는 주관적인 안전성 인식, 평가자 간 이견, 안전성 판단에 영향을 미치는 다차원적 인구통계학적 요인을 깊이 있게 분석할 수 있게 하며, 기존의 '황금 기준 레이블' 개념을 도전하고, 포괄적이고 편향 인식이 있는 모델 평가 기반을 제공한다.

ABSTRACT

Machine learning approaches often require training and evaluation datasets with a clear separation between positive and negative examples. This risks simplifying and even obscuring the inherent subjectivity present in many tasks. Preserving such variance in content and diversity in datasets is often expensive and laborious. This is especially troubling when building safety datasets for conversational AI systems, as safety is both socially and culturally situated. To demonstrate this crucial aspect of conversational AI safety, and to facilitate in-depth model performance analyses, we introduce the DICES (Diversity In Conversational AI Evaluation for Safety) dataset that contains fine-grained demographic information about raters, high replication of ratings per item to ensure statistical power for analyses, and encodes rater votes as distributions across different demographics to allow for in-depth explorations of different aggregation strategies. In short, the DICES dataset enables the observation and measurement of variance, ambiguity, and diversity in the context of conversational AI safety. We also illustrate how the dataset offers a basis for establishing metrics to show how raters' ratings can intersects with demographic categories such as racial/ethnic groups, age groups, and genders. The goal of DICES is to be used as a shared resource and benchmark that respects diverse perspectives during safety evaluation of conversational AI systems.

연구 동기 및 목표

  • 안전성이 문화적·사회적으로 뿌리를 두고 있음에도 불구하고 대화형 AI 안전성 평가에서 다양하고 대표적인 시각의 부족을 해결하기 위해.
  • 안전성에 대한 평가자 의견의 다양성, 모호성, 변동성을 반영하여 단순한 '기본 진실' 레이블을 넘어서기 위해.
  • 민족/인종, 연령, 성별 등의 인구통계학적 요인이 안전성 판단에 어떻게 영향을 미치는지 분석할 수 있도록 통계적 힘과 깊이 있는 분석을 지원하는 벤치마크를 제공하기 위해.
  • 단일 기준을 강요하는 것이 아니라 다양한 안전성 개념을 반영할 수 있는 피팅 및 평가 방법의 개발을 가능하게 하기 위해.
  • 높은 평가자 이견이 존재함을 입증함으로써 기존의 '황금 기준' 개념을 도전하고, 데이터셋 품질과 모델 훈련에 미치는 영향을 제기하기 위해.

제안 방법

  • 성별, 연령, 민족/인종 집단에 따라 균형 잡힌 평가자 풀을 구성하여 다양한 배경의 인구통계학적 표현을 확보한다.
  • DICES-990(990개 대화)와 DICES-350(350개 대화)의 각 대화는 70~120건의 평가를 확보하여 높은 통계적 힘과 변동성 추정을 위한 재표본 추출을 가능하게 한다.
  • 해로움, 편향, 오락성, 정치, 정책 위반 등 다섯 가지 안전성 범주에 걸쳐 평가하며, 혐오 발언과 같은 특정 유형에 대한 하위 평가도 포함한다.
  • 평가자 투표는 인구통계학적 집단별 분포로 인코딩되어 다수 수준의 베이지안 모델링을 통해 이견과 다중 축 인구통계학적 영향을 분석할 수 있다.
  • 안전성, 해로움, 해로움 정도에 대한 전문가 주석이 포함되어 있어 커뮤니티 평가와의 비교를 가능하게 한다.
  • 시간적 및 행동 데이터를 활용해 이견 지표와 평가자 행동을 분석하여 이질적 사례와 노이즈를 탐지할 수 있다.
Figure 2: Screenshot of the raters’ user interface for the Safety Annotation Task: illustrates the annotation category for policy violations . The left panel presents the conversation; raters assess the last conversational turn (highlighted). The right panel presents two policy related sub-questions
Figure 2: Screenshot of the raters’ user interface for the Safety Annotation Task: illustrates the annotation category for policy violations . The left panel presents the conversation; raters assess the last conversational turn (highlighted). The right panel presents two policy related sub-questions

실험 결과

연구 질문

  • RQ1다양한 인구통계학적 집단(예: 인종, 연령, 성별)은 대화형 AI 출력물의 안전성에 대해 어떻게 다른 인식을 가지는가?
  • RQ2평가자 다양성이 안전성 판단의 이견에 얼마나 큰 영향을 미치며, 이를 통계적으로 어떻게 모델링할 수 있는가?
  • RQ3다중 축 인구통계학적 정체성(예: 젊은 블랙 여성)은 단일 축 기반 그룹에 비해 안전성 인식에 어떻게 영향을 미치는가?
  • RQ4평가자 간 이견이 높을 경우 전문가 주석과 커뮤니티 평가를 의미 있게 비교할 수 있는가?
  • RQ5높은 평가자 복제 수가 변동성 추정과 모델 평가의 강건성 향상에 어떤 영향을 미치는가?

주요 결과

  • DICES 데이터셋은 1,340개의 대화에서 약 300명의 평가자로부터 250만 건 이상의 평가를 확보하였으며, 1회 대화당 70~120건의 평가를 확보하여 일반적인 3~5명의 평가 기준을 크게 초월한다.
  • 평가자 간 이견이 높아 안전성 평가에 상당한 주관성이 있음을 시사하며, 안전성에 대한 단일 '황금 기준'의 실현 가능성에 의문을 제기한다.
  • 모든 인구통계학적 집단이 안전성 인식에 동일한 영향을 미치지는 않으며, 일부 하위 집단은 평가 패턴에서 더 강하거나 일관성 있는 경향을 보인다.
  • 다중 축 정체성(예: 젊은 블랙 여성)은 단일 축 기반 분석에서 놓칠 수 있는 특이한 안전성 인식 패턴을 드러내며, 이는 단일 축 분석이 중요한 통찰을 놓칠 수 있음을 시사한다.
  • 높은 복제 수 덕분에 안정적인 재표본 추출과 변동성의 더 나은 추정이 가능하여, 안전성 평가 연구에서 더 신뢰할 수 있는 통계적 추론을 지원한다.
  • 전문가 평가와 커뮤니티 평가 간 이격이 관찰되어, 단일 전문가 공감대에 의존하기보다는 다양한 시각을 통합하는 방법의 필요성을 강조한다.
Figure 3: Breakdown of topics and degree of harm for DICES-350. Percentages of conversations per topic (left) and number of conversations per degree of harm (right).
Figure 3: Breakdown of topics and degree of harm for DICES-350. Percentages of conversations per topic (left) and number of conversations per degree of harm (right).

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.