Skip to main content
QUICK REVIEW

[论文解读] Towards Measuring Fairness in AI: the Casual Conversations Dataset

Caner Hazırbaş, Joanna Bitton|arXiv (Cornell University)|Apr 6, 2021
Building Energy and Comfort Optimization被引用 4
一句话总结

本文介绍了 Casual Conversations 数据集,这是一个多样化的、由人类标注的视频集合,包含 3,011 名受试者,其自报年龄、性别、Fitzpatrick 皮肤类型及光照条件,用于评估人工智能的公平性。研究发现,顶级深度伪造检测模型和最先进的年龄/性别模型在深色皮肤和低光照条件下性能显著下降,凸显了当前人工智能系统中的关键偏见。

ABSTRACT

This paper introduces a novel dataset to help researchers evaluate their computer vision and audio models for accuracy across a diverse set of age, genders, apparent skin tones and ambient lighting conditions. Our dataset is composed of 3,011 subjects and contains over 45,000 videos, with an average of 15 videos per person. The videos were recorded in multiple U.S. states with a diverse set of adults in various age, gender and apparent skin tone groups. A key feature is that each subject agreed to participate for their likenesses to be used. Additionally, our age and gender annotations are provided by the subjects themselves. A group of trained annotators labeled the subjects' apparent skin tone using the Fitzpatrick skin type scale. Moreover, annotations for videos recorded in low ambient lighting are also provided. As an application to measure robustness of predictions across certain attributes, we provide a comprehensive study on the top five winners of the DeepFake Detection Challenge (DFDC). Experimental evaluation shows that the winning models are less performant on some specific groups of people, such as subjects with darker skin tones and thus may not generalize to all people. In addition, we also evaluate the state-of-the-art apparent age and gender classification methods. Our experiments provides a thorough analysis on these models in terms of fair treatment of people from various backgrounds.

研究动机与目标

  • 通过创建一个多样化、具有代表性的数据集,评估不同人口统计和环境属性下的公平性,以解决人工智能模型中的算法偏见。
  • 为在不同条件(如皮肤色调和环境光照)下计算机视觉和音频模型的鲁棒性提供基准。
  • 评估来自深度伪造检测挑战赛(DFDC)的顶尖模型在不同人口群体中的真实世界泛化能力。
  • 通过包含自报年龄和性别的数据集,减少潜在的标注偏见,推动更包容和负责任的人工智能开发。

提出的方法

  • 该数据集包含来自美国不同州的 3,011 名参与者,每人贡献约 15 段非正式对话视频。
  • 参与者自报其年龄和性别,以减少人口统计属性中潜在的标注偏差。
  • 由训练有素的标注员使用 Fitzpatrick 皮肤类型量表对表观皮肤色调进行标注,实现对六种皮肤类型类别的分析。
  • 对环境光照条件(明亮 vs. 昏暗)进行标注,以支持在低光照场景下的评估。
  • 使用 DLIB 进行人脸检测,并对每段视频的 100 个采样人脸区域进行模型预测聚合,以提高鲁棒性。
  • 评估聚焦于深度伪造检测模型和最先进的年龄/性别分类模型,采用不同公平性类别下的精确率、对数损失和 ROC 曲线分析。

实验结果

研究问题

  • RQ1来自 DFDC 的顶尖深度伪造检测模型在不同皮肤色调和光照条件下表现如何?
  • RQ2最先进的表观年龄和性别分类模型在多大程度上对浅色皮肤或特定年龄/性别群体表现出偏见?
  • RQ3与第三方标注相比,自报年龄和性别标注是否能减少人口统计属性数据集中的偏见?
  • RQ4环境光照如何影响面部属性识别和深度伪造检测模型的性能?
  • RQ5在 Fitzpatrick 皮肤类型中,模型性能的公平性差距如何,特别是对深色皮肤(类型 V 和 VI)?

主要发现

  • 包括冠军 Selim Seferbekov [8] 在内的顶级 DFDC 模型在深色皮肤(类型 VI)上的表现显著下降,精确率相比浅色皮肤类型下降超过 20%。
  • 表观性别分类模型在类型 VI 皮肤上的平均精确率仅为 53.78%,而类型 I 上为 83.33%,表明存在显著的公平性差距。
  • 表现最佳的年龄与性别模型 LightFace [35] 在深色皮肤上的性能仍下降超过 20%,表明当前最先进系统中仍存在持续的偏见。
  • 所有评估模型在低环境光照条件下录制的视频中准确率均下降,昏暗条件下精确率最高下降 10%。
  • 如 The Medics [12] 和 Eighteen Years Old [11] 等模型在各类别中表现出不一致的性能,年龄、性别或光照条件之间无明显平衡。
  • 该数据集表明,基于不平衡数据集训练的模型无法在多样化人群中泛化,尤其对深色皮肤人群等代表性不足群体表现更差。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。