Skip to main content
QUICK REVIEW

[论文解读] Consensus and Subjectivity of Skin Tone Annotation for ML Fairness

Candice Schumann, Gbolahan O. Olanubi|arXiv (Cornell University)|May 16, 2023
Infection Control and Ventilation被引用 14
一句话总结

本论文研究 Monk Skin Tone (MST) 量表的皮肤色调注释如何因注释者类型和地理分布而异,提出用于训练和评估的 MST-E 数据集,并为公平性研究中的多样化、可重复的注释提供最佳实践指南。

ABSTRACT

Understanding different human attributes and how they affect model behavior may become a standard need for all model creation and usage, from traditional computer vision tasks to the newest multimodal generative AI systems. In computer vision specifically, we have relied on datasets augmented with perceived attribute signals (e.g., gender presentation, skin tone, and age) and benchmarks enabled by these datasets. Typically labels for these tasks come from human annotators. However, annotating attribute signals, especially skin tone, is a difficult and subjective task. Perceived skin tone is affected by technical factors, like lighting conditions, and social factors that shape an annotator's lived experience. This paper examines the subjectivity of skin tone annotation through a series of annotation experiments using the Monk Skin Tone (MST) scale, a small pool of professional photographers, and a much larger pool of trained crowdsourced annotators. Along with this study we release the Monk Skin Tone Examples (MST-E) dataset, containing 1515 images and 31 videos spread across the full MST scale. MST-E is designed to help train human annotators to annotate MST effectively. Our study shows that annotators can reliably annotate skin tone in a way that aligns with an expert in the MST scale, even under challenging environmental conditions. We also find evidence that annotators from different geographic regions rely on different mental models of MST categories resulting in annotations that systematically vary across regions. Given this, we advise practitioners to use a diverse set of annotators and a higher replication count for each image when annotating skin tone for fairness research.

研究动机与目标

  • 评估注释者类型(专家 vs 众包)如何影响 MST 皮肤色调注释。
  • 检验地理区域对 MST 注释和模型注释者行为的影响。
  • 提供数据集和培训资源以提高 MST 注释的一致性。
  • 就设计皮肤色调注释任务在公平性研究中的实际建议。

提出的方法

  • 引入 Monk Skin Tone (MST) 量表以及包含 1515 张图像和 31 段视频、覆盖 10 个 MST 点的 MST-E 数据集。
  • 进行两项注释实验:一项小型的专家摄影师研究,和一项跨五个区域的大型众包注释者研究。
  • 将注释者的中位数注释与由 MST 量表创建者 Dr. Ellis Monk 提供的金标准 oracle 进行比较。
  • 使用组内相关系数 (ICC) 测量评注者之间的一致性,并使用 1 点差距和平均中位距离度量来评估与 oracle 的一致性。
  • 将注释实验扩展到野外的 Open Images 数据,以测试注释者行为的泛化性。

实验结果

研究问题

  • RQ1谁在 MST 量表上可以可靠地进行皮肤色调注释(专家 vs 众包),以及在何种条件下?
  • RQ2注释者的地理区域是否会影响 MST 注释,应该如何进行管理?
  • RQ3经过培训的注释者是否能够在不同光照条件下实现接近 MST 创作者意图的共识?
  • RQ4哪些实际的注释设计建议可以提高可靠性和公平性分析的效果?

主要发现

  • 无论来自专家还是众包群体,注释者在不同光照条件下都能可靠地对 MST 进行注释,并与 MST 创作者的意图保持一致。
  • 区域差异较大:印度摄影师在同一对象上倾向标注较浅的 MST,而美国摄影师则倾向标注较深的 MST。
  • 在专家和众包研究中,相当大比例的共识注释与 oracle 的距离在 1 MST 点内(印度为 88.9%,美国为 83.4%,专家组;众包组显示出高 ICC)。
  • 跨五个区域的众包注释者显示出强烈的评注者间一致性(对每个主体 ICC 0.86–0.94;对 Golden Images ICC 0.90–0.96),与 oracle 的距离平均仍在 1 点以下(0.78–0.84)。
  • MST-E 数据集支持在整个 MST 量表及不同光照条件下对公平性进行注释者和模型的训练与评估;多样化区域的注释者群体与 MST 量表创建者的意图保持一致。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。