Skip to main content
QUICK REVIEW

[Paper Review] Consensus and Subjectivity of Skin Tone Annotation for ML Fairness

Candice Schumann, Gbolahan O. Olanubi|arXiv (Cornell University)|May 16, 2023
Infection Control and Ventilation14 citations
TL;DR

The paper investigates how skin tone annotations on the Monk Skin Tone (MST) scale vary by annotator type and geography, introduces the MST-E dataset for training and evaluation, and provides best-practice guidance for diverse, replicated annotations in fairness research.

ABSTRACT

Understanding different human attributes and how they affect model behavior may become a standard need for all model creation and usage, from traditional computer vision tasks to the newest multimodal generative AI systems. In computer vision specifically, we have relied on datasets augmented with perceived attribute signals (e.g., gender presentation, skin tone, and age) and benchmarks enabled by these datasets. Typically labels for these tasks come from human annotators. However, annotating attribute signals, especially skin tone, is a difficult and subjective task. Perceived skin tone is affected by technical factors, like lighting conditions, and social factors that shape an annotator's lived experience. This paper examines the subjectivity of skin tone annotation through a series of annotation experiments using the Monk Skin Tone (MST) scale, a small pool of professional photographers, and a much larger pool of trained crowdsourced annotators. Along with this study we release the Monk Skin Tone Examples (MST-E) dataset, containing 1515 images and 31 videos spread across the full MST scale. MST-E is designed to help train human annotators to annotate MST effectively. Our study shows that annotators can reliably annotate skin tone in a way that aligns with an expert in the MST scale, even under challenging environmental conditions. We also find evidence that annotators from different geographic regions rely on different mental models of MST categories resulting in annotations that systematically vary across regions. Given this, we advise practitioners to use a diverse set of annotators and a higher replication count for each image when annotating skin tone for fairness research.

Motivation & Objective

  • Assess how annotator type (experts vs crowdsourced) influences MST skin tone annotations.
  • Examine the impact of geographic region on MST annotations and model annotator behavior.
  • Provide a dataset and training resources to improve consistency in MST annotations.
  • Offer practical recommendations for designing skin tone annotation tasks in fairness research.

Proposed method

  • Introduce the Monk Skin Tone (MST) scale and the MST-E dataset containing 1515 images and 31 videos across 10 MST points.
  • Conduct two annotation experiments: a small expert-photographer study and a larger crowdsourced annotator study across five regions.
  • Compare annotator median annotations to a gold standard oracle provided by MST scale creator Dr. Ellis Monk.
  • Measure inter-rater reliability using intraclass correlation (ICC) and assess agreement with the oracle using 1-point discrepancy and average median distance metrics.
  • Extend annotation experiments to in-the-wild Open Images data to test generalizability of annotator behavior.

Experimental results

Research questions

  • RQ1Who can reliably annotate skin tone on the MST scale (experts vs crowdsourced) and under what conditions?
  • RQ2Does annotator geographic region affect MST annotations and how should this be managed?
  • RQ3Can trained annotators achieve consensus close to the MST creator's intent across lighting conditions?
  • RQ4What practical annotation design recommendations improve reliability and fairness analyses?

Key findings

  • Annotators from both experts and crowdsourced pools reliably annotate MST, aligning with the MST creator’s intent under varied lighting conditions.
  • Higher regional differences were observed: Indian photographers tended to label lighter MSTs while US photographers tended toward darker MSTs for the same subjects.
  • In both expert and crowdsource studies, a large majority of consensus annotations were within 1 MST point of the oracle (88.9% in India, 83.4% in the US for experts; high ICCs in crowdsourced groups).
  • Crowdsourced annotators across five regions showed strong inter-rater reliability (ICC 0.86–0.94 per subject; 0.90–0.96 for Golden Images), and distance to oracle remained under 1 point on average (0.78–0.84).
  • The MST-E dataset supports training and evaluating annotators and models for fairness across the full MST scale and varied lighting; diverse regional annotator pools yield annotations aligned with the MST scale creator’s intent.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.