Skip to main content
QUICK REVIEW

[论文解读] Analysis of Manual and Automated Skin Tone Assignments for Face Recognition Applications

K. S. Krishnapriya, Michael C. King|arXiv (Cornell University)|Apr 29, 2021
Face recognition and analysis参考文献 34被引用 7
一句话总结

本文评估了使用Fitzpatrick量表和个体类型角(ITA)进行人脸虹膜识别中人工与自动肤色分配的效果。研究发现,人工评估者在肤色标签标注上存在显著不一致性,但在色彩校正后,基于ITA的自动估计方法与人工共识的吻合度达到96%或以上,优于人工评估者间的一致性,为偏见分析提供了可扩展、可复现的标签方法。

ABSTRACT

News reports have suggested that darker skin tone causes an increase in face recognition errors. The Fitzpatrick scale is widely used in dermatology to classify sensitivity to sun exposure and skin tone. In this paper, we analyze a set of manual Fitzpatrick skin type assignments and also employ the individual typology angle to automatically estimate the skin tone from face images. The set of manual skin tone rating experiments shows that there are inconsistencies between human raters that are difficult to eliminate. Efforts to automate skin tone rating suggest that it is particularly challenging on images collected without a calibration object in the scene. However, after the color-correction, the level of agreement between automated and manual approaches is found to be 96% or better for the MORPH images. To our knowledge, this is the first work to: (a) examine the consistency of manual skin tone ratings across observers, (b) document that there is substantial variation in the rating of the same image by different observers even when exemplar images are given for guidance and all images are color-corrected, and (c) compare manual versus automated skin tone ratings.

研究动机与目标

  • 评估使用Fitzpatrick量表对人脸图像进行肤色评级的人工一致性。
  • 评估图像色彩差异与光照对人工肤色评级可靠性的影响。
  • 开发并验证一种基于个体类型角(ITA)的自动化肤色估计方法。
  • 比较基于自动化ITA的肤色分配与人工共识评级在准确性和一致性方面的表现。
  • 为面部识别偏见研究提供一种可扩展、可复现的肤色标签方法。

提出的方法

  • 在MORPH数据集的色彩校正人脸图像上,使用Fitzpatrick量表开展人工肤色评级实验。
  • 向评估者提供参考色卡和示例图像,以提高人工标注的一致性。
  • 实现了一套基于人脸图像中个体类型角(ITA)测量的自动化肤色估计流程。
  • 通过色彩校正减少光照与传感器差异的影响。
  • 定义了定制化的ITA阈值范围,将连续的ITA值映射为离散的Fitzpatrick类肤色类别(I–VI)。
  • 使用类别相似性度量指标量化人工共识评级与自动化ITA输出之间的一致性。

实验结果

研究问题

  • RQ1即使在提供指导和色彩校正的情况下,人工评估者在为人脸图像分配Fitzpatrick肤色类别时是否具有一致性?
  • RQ2图像光照与传感器差异在多大程度上影响了人工肤色评级的可靠性?
  • RQ3基于自动化ITA的方法与人工共识评级在肤色分类上的一致性如何?
  • RQ4自动化肤色分配能否达到与人工评估者间一致性的水平?
  • RQ5色彩校正对减少人工与自动化方法在肤色标签上的变异性有何影响?

主要发现

  • 即使使用示例图像和色彩校正图像,人工评估者在肤色评级上仍表现出显著不一致性,当不允许容忍度时,评估者间的一致性低于90%。
  • 经过色彩校正后,人工共识评级与基于ITA的自动化肤色分配之间的一致性达到96%或以上,表明其具有极高的可靠性。
  • 自动化ITA与人工共识之间的一致性水平与任意两名人工评估者之间的评估一致性相当,验证了自动化方法的一致性。
  • 本研究发现,光照与传感器差异是人工与自动化肤色估计中误分类的主要原因,尤其在非受控环境中更为显著。
  • 尽管经过色彩校正和示例引导,人工评估者之间的分歧依然存在,表明在类别化肤色标签中存在固有的主观性。
  • 基于ITA的自动化方法为大规模人脸虹膜识别偏见研究提供了一种可扩展、可复现且一致的替代人工标注的方法。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。