[论文解读] One Label, One Billion Faces: Usage and Consistency of Racial Categories in Computer Vision
本文通过训练分类器以学习标注人脸数据集中隐含的种族体系,研究了计算机视觉数据集中种族类别的跨数据集一致性。研究发现,种族类别在不同数据集中高度不一致,常常编码种族刻板印象并排除非刻板印象的族裔群体,从而损害人工智能中的公平性度量和多样性声明。
Computer vision is widely deployed, has highly visible, society altering applications, and documented problems with bias and representation. Datasets are critical for benchmarking progress in fair computer vision, and often employ broad racial categories as population groups for measuring group fairness. Similarly, diversity is often measured in computer vision datasets by ascribing and counting categorical race labels. However, racial categories are ill-defined, unstable temporally and geographically, and have a problematic history of scientific use. Although the racial categories used across datasets are superficially similar, the complexity of human race perception suggests the racial system encoded by one dataset may be substantially inconsistent with another. Using the insight that a classifier can learn the racial system encoded by a dataset, we conduct an empirical study of computer vision datasets supplying categorical race labels for face images to determine the cross-dataset consistency and generalization of racial categories. We find that each dataset encodes a substantially unique racial system, despite nominally equivalent racial categories, and some racial categories are systemically less consistent than others across datasets. We find evidence that racial categories encode stereotypes, and exclude ethnic groups from categories on the basis of nonconformity to stereotypes. Representing a billion humans under one racial category may obscure disparities and create new ones by encoding stereotypes of racial systems. The difficulty of adequately converting the abstract concept of race into a tool for measuring fairness underscores the need for a method more flexible and culturally aware than racial categories.
研究动机与目标
- 调查计算机视觉数据集中使用的种族类别的跨数据集一致性。
- 评估种族类别是否编码了种族刻板印象并排除了不符合规范的族裔群体。
- 评估种族标注在不同数据集之间的一般化和可靠性。
- 强调在人工智能系统中将种族视为固定、分类标签的风险。
- 质疑将种族类别用作机器学习中公平性和多样性代理指标的有效性。
提出的方法
- 在四个主要的公平计算机视觉数据集上训练多个图像分类器,这些数据集包含分类的种族标签。
- 使用分类器在不同数据集之间的共识作为衡量种族类别边界一致性的代理指标。
- 评估个体在不同数据集中被分配到种族类别的连续性,特别是针对非刻板印象族裔群体。
- 收集了一个小规模人工标注的数据集,以测试族裔群体与数据集中编码的种族刻板印象的符合程度。
- 分析来自东非(例如埃塞俄比亚人)与西非/中非(例如冈比亚人)的个体在标注上的差异,以检测地理和文化偏见。
- 比较不同数据集之间的分类器预测结果,以检测系统性刻板印象和排除模式。
实验结果
研究问题
- RQ1尽管使用了名义上相同的标签,不同计算机视觉数据集中种族类别的分配有多一致?
- RQ2数据集中的种族类别在多大程度上编码了种族刻板印象,而非社会文化或族裔身份?
- RQ3非刻板印象族裔群体(例如东非人)在不同数据集中的标注一致性如何?
- RQ4不一致的种族标注对人工智能中公平性评估和多样性测量有何影响?
- RQ5数据集收集的文化和地理背景如何影响计算机视觉中种族类别的构建?
主要发现
- 每个数据集都编码了一个显著独特的种族体系,即使使用了相同的名义种族类别标签。
- 种族类别在标注上表现出系统性不一致,某些群体(例如东非人)在不同数据集中被不一致地标注为“黑人”。
- 在不同数据集上训练的分类器对于那些与某一类别主导身体特征不相符的个体的种族判断存在显著分歧。
- 研究发现,埃塞俄比亚人被标注为“黑人”的一致性低于冈比亚人,表明数据集构建中存在地理和文化偏见。
- 数据集中的种族类别常常排除或错误描述不符合主流身体刻板印象的族裔群体。
- 通过数据集将种族实体化,可能加剧历史和文化偏见,尤其是在数据集被跨人工智能系统重复使用和传播时。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。