[论文解读] Can you tell where in India I am from? Comparing humans and computers on fine-grained race face classification
本研究提出了一项细粒度的种族分类任务,旨在通过一种新型数据集区分北印度人与南印度人脸。该数据集包含1,647张多样化人脸,经人工标注并记录了129名受试者的性能数据。尽管机器在整体准确率上与人类相当(64%),但其错误模式与人类系统性不同,揭示出口型是最具区分力的面部特征——这一发现通过遮挡实验得到证实,即当遮挡口部时,人类表现下降最明显。
Faces form the basis for a rich variety of judgments in humans, yet the underlying features remain poorly understood. Although fine-grained distinctions within a race might more strongly constrain possible facial features used by humans than in case of coarse categories such as race or gender, such fine grained distinctions are relatively less studied. Fine-grained race classification is also interesting because even humans may not be perfectly accurate on these tasks. This allows us to compare errors made by humans and machines, in contrast to standard object detection tasks where human performance is nearly perfect. We have developed a novel face database of close to 1650 diverse Indian faces labeled for fine-grained race (South vs North India) as well as for age, weight, height and gender. We then asked close to 130 human subjects who were instructed to categorize each face as belonging toa Northern or Southern state in India. We then compared human performance on this task with that of computational models trained on the ground-truth labels. Our main results are as follows: (1) Humans are highly consistent (average accuracy: 63.6%), with some faces being consistently classified with > 90% accuracy and others consistently misclassified with < 30% accuracy; (2) Models trained on ground-truth labels showed slightly worse performance (average accuracy: 62%) but showed higher accuracy (72.2%) on faces classified with > 80% accuracy by humans. This was true for models trained on simple spatial and intensity measurements extracted from faces as well as deep neural networks trained on race or gender classification; (3) Using overcomplete banks of features derived from each face part, we found that mouth shape was the single largest contributor towards fine-grained race classification, whereas distances between face parts was the strongest predictor of gender.
研究动机与目标
- 调查人类在一项困难的细粒度种族分类任务中对北印度人与南印度人脸的区分表现。
- 识别人类在进行此类细粒度区分时所依赖的具体面部特征,以克服粗粒度种族分类的局限性。
- 将机器学习模型在相同任务上的表现与错误模式与人类表现进行对比。
- 通过行为遮挡实验验证特定面部区域的重要性。
- 揭示人类与机器在面部识别中表征学习的系统性差异。
提出的方法
- 构建了一个包含1,647张多样化印度人脸的新数据集,标注了其北印度或南印度起源,并记录了129名人类受试者的性能数据。
- 在从各个面部区域(眼睛、鼻子、嘴巴、轮廓)提取的过完备特征集上训练线性分类器,以识别具有区分力的特征。
- 对24名人类受试者开展受控的行为实验,在不同条件下分别遮挡眼睛、鼻子或嘴巴,以评估各特征的重要性。
- 使用Wilcoxon符号秩检验比较不同遮挡条件下分类准确率与反应时间的差异。
- 在数据集上评估多种机器学习模型,并将其错误模式与人类表现进行对比。
- 每种遮挡条件(未遮挡、眼遮挡、鼻遮挡、口遮挡)选取217张人脸,确保各条件下基线准确率相近(约69%)。
实验结果
研究问题
- RQ1人类在区分印度人脸的细粒度区域种族差异时,使用了哪些面部特征?
- RQ2与人类相比,机器学习模型在该细粒度分类任务中的表现如何?
- RQ3在该困难分类任务中,机器与人类的错误模式是否系统性不同?
- RQ4遮挡特定面部区域是否会降低人类的分类准确率?如果是,哪个区域影响最大?
- RQ5基于面部特征的计算建模能否预测人类面部识别中各面部部分的相对重要性?
主要发现
- 在细粒度的北印度人与南印度人脸分类任务中,人类在未遮挡人脸上的平均准确率为65.8%。
- 机器达到了与人类相当的整体准确率(64%),但其在不同人脸上的错误模式与人类存在定性差异。
- 遮挡口部对人类分类准确率的影响最大(59.8%),显著低于遮挡眼睛(63.6%)或鼻子(61.1%),与未遮挡条件相比p < 0.0005。
- 口部遮挡条件下的反应时间显著更长(与未遮挡相比p < 0.0005),证实认知负荷增加且对口型的依赖性更强。
- 基于局部特征训练的线性分类器识别出口型为种族分类中最具区分力的特征,支持行为实验结果。
- 机器与人类在错误模式上的系统性差异表明,人类使用的是标准机器表征无法捕捉的、非可解释的特征。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。