Skip to main content
QUICK REVIEW

[논문 리뷰] Consensus and Subjectivity of Skin Tone Annotation for ML Fairness

Candice Schumann, Gbolahan O. Olanubi|arXiv (Cornell University)|2023. 05. 16.
Infection Control and Ventilation인용 수 14
한 줄 요약

본 논문은 Monk Skin Tone (MST) 척도상의 피부 톤 주석이 주석자 유형 및 지리적 위치에 따라 어떻게 달라지는지 조사하고, 훈련 및 평가를 위한 MST-E 데이터셋을 도입하며, 공정성 연구에서 다양한 재현 가능한 주석을 위한 모범 사례 지침을 제시한다.

ABSTRACT

Understanding different human attributes and how they affect model behavior may become a standard need for all model creation and usage, from traditional computer vision tasks to the newest multimodal generative AI systems. In computer vision specifically, we have relied on datasets augmented with perceived attribute signals (e.g., gender presentation, skin tone, and age) and benchmarks enabled by these datasets. Typically labels for these tasks come from human annotators. However, annotating attribute signals, especially skin tone, is a difficult and subjective task. Perceived skin tone is affected by technical factors, like lighting conditions, and social factors that shape an annotator's lived experience. This paper examines the subjectivity of skin tone annotation through a series of annotation experiments using the Monk Skin Tone (MST) scale, a small pool of professional photographers, and a much larger pool of trained crowdsourced annotators. Along with this study we release the Monk Skin Tone Examples (MST-E) dataset, containing 1515 images and 31 videos spread across the full MST scale. MST-E is designed to help train human annotators to annotate MST effectively. Our study shows that annotators can reliably annotate skin tone in a way that aligns with an expert in the MST scale, even under challenging environmental conditions. We also find evidence that annotators from different geographic regions rely on different mental models of MST categories resulting in annotations that systematically vary across regions. Given this, we advise practitioners to use a diverse set of annotators and a higher replication count for each image when annotating skin tone for fairness research.

연구 동기 및 목표

  • 주석자 유형(전문가 대 crowdsourced)이 MST 피부 톤 주석에 어떤 영향을 미치는지 평가한다.
  • 지리적 지역이 MST 주석과 모델 주석자 행동에 미치는 영향을 검토한다.
  • MST 주석의 일관성을 개선하기 위한 데이터셋과 교육 자원을 제공한다.
  • 공정성 연구에서 피부 톤 주석 작업을 설계하기 위한 실용적 권고안을 제시한다.

제안 방법

  • Monk Skin Tone (MST) 척도와 10 MST 포인트에 걸친 1515 이미지 및 31 비디오를 포함하는 MST-E 데이터셋을 도입한다.
  • 두 가지 주석 실험을 수행한다: 소규모 전문가-사진가 연구와 다섯 지역에 걸친 더 큰 crowdsourced 주석자 연구.
  • 주석자 중앙값 주석을 MST 척도 창시자인 Dr. Ellis Monk가 제공한 골드 스탠다드 오라클과 비교한다.
  • 관찰자 간 신뢰도(inter-rater reliability)를 intraclass correlation (ICC)로 측정하고 1포인트 차이 및 평균 중앙값 거리 지표를 사용하여 오라클과의 합의 여부를 평가한다.
  • 주석 실험을 실제 환경의 Open Images 데이터로 확장하여 주석자 행동의 일반화 가능성을 테스트한다.

실험 결과

연구 질문

  • RQ1누가 MST 척도에서 피부 톤을 신뢰할 수 있게 주석할 수 있는가(전문가 대 crowdsourced)와 어떤 조건에서 그런가?
  • RQ2주석자의 지리적 지역이 MST 주석에 영향을 미치는가, 그리고 이를 어떻게 관리해야 하는가?
  • RQ3조명 조건에 걸쳐 MST 창시자의 의도에 근접한 합의에 도달하는 교육된 주석자는 있을 수 있는가?
  • RQ4신뢰성과 공정성 분석을 개선하는 실용적인 주석 설계 권고 사항은 무엇인가?

주요 결과

  • 전문가와 crowdsourced 양쪽 주석자 풀의 주석자들이 다양한 조명 조건에서 MST 창시자의 의도와 일치하는 신뢰할 수 있는 주석을 내릴 수 있다.
  • 지역 간 차이가 더 크게 나타났는데: 인도 사진가는 동일한 피사체에 대해 더 밝은 MST로 라벨하는 경향이 있었고 미국 사진가는 더 어두운 MST로 라벨하는 경향이 있었다.
  • 전문가와 crowdsourcing 연구 모두에서 합의 주석의 대다수가 오라클과 1 MST 포인트 이내에 위치했다(인도 88.9%, 미국 83.4%로 전문가; crowdsourced 그룹은 높은 ICC).
  • 다섯 지역의 crowdsourced 주석자들은 높은 관찰자 간 신뢰도(ICCs 0.86–0.94 각 피사체; Golden Images의 경우 0.90–0.96)를 보였고, 오라클과의 거리는 평균적으로 1포인트 미만으로 유지되었다(0.78–0.84).
  • MST-E 데이터셋은 전체 MST 스케일과 다양한 조명에서 공정성을 위한 주석자 및 모델의 훈련과 평가를 지원하며, 다양한 지역 주석자 풀은 MST 창시자의 의도와 일치하는 주석을 도출한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.