[论文解读] An Examination of Fairness of AI Models for Deepfake Detection
本研究通过分析数据集和模型在种族与性别上的偏差,探究深度伪造检测人工智能模型中的公平性问题。研究发现,广泛使用的数据集如FaceForensics++严重偏向白人女性,且深度伪造生成方法会产生'不规则'的面部替换,导致检测器将虚假性与特定种族特征关联,从而造成某些种族子群体的错误率最高上升10.7%,尤其是亚裔女性。
Recent studies have demonstrated that deep learning models can discriminate based on protected classes like race and gender. In this work, we evaluate bias present in deepfake datasets and detection models across protected subgroups. Using facial datasets balanced by race and gender, we examine three popular deepfake detectors and find large disparities in predictive performances across races, with up to 10.7% difference in error rate between subgroups. A closer look reveals that the widely used FaceForensics++ dataset is overwhelmingly composed of Caucasian subjects, with the majority being female Caucasians. Our investigation of the racial distribution of deepfakes reveals that the methods used to create deepfakes as positive training signals tend to produce "irregular" faces - when a person's face is swapped onto another person of a different race or gender. This causes detectors to learn spurious correlations between the foreground faces and fakeness. Moreover, when detectors are trained with the Blended Image (BI) dataset from Face X-Rays, we find that those detectors develop systematic discrimination towards certain racial subgroups, primarily female Asians.
研究动机与目标
- 调查深度伪造检测模型在种族和性别等受保护属性上的公平性。
- 分析FaceForensics++和Blended Images等广泛使用的深度伪造数据集中存在的表征偏差。
- 研究深度伪造的数据生成方法如何在面部特征与虚假性之间产生虚假相关性。
- 评估深度伪造检测器在交叉种族与性别子群体中的预测性能差异。
- 倡导在合成媒体的人工智能系统中采用更具代表性的数据和交叉性审计。
提出的方法
- 标注者对FaceForensics++和Blended Images中的1,000多段真实视频按性别和种族进行标注,达到75.93%的评分者间一致性。
- 本研究将面部替换分类为'规则'(同种族/同性别)或'不规则'(不同种族/不同性别),以评估训练数据中的偏差。
- 构建了一个性别与种族均衡的面部数据集,以评估检测器的公平性。
- 在该均衡数据集上测试了三种流行的深度伪造检测器,以衡量性能差异。
- 研究人员分析了FF++和BI数据集中虚假样本的前景和背景人脸分布,以识别偏差的训练信号。
- 通过统计分析比较了交叉子群体间的错误率,揭示了假阳性率中的系统性差异。
实验结果
研究问题
- RQ1FaceForensics++数据集在现实世界深度伪造检测场景中,对多样化种族与性别群体的代表性如何?
- RQ2深度伪造的数据生成方法在面部特征与虚假性之间如何导致虚假相关性?
- RQ3深度伪造检测器在交叉种族与性别子群体中的性能差异是什么?
- RQ4在Blended Images数据集上进行训练如何影响公平性,特别是对亚裔女性的影响?
- RQ5训练数据中'不规则'的面部替换(跨种族或跨性别)在多大程度上导致检测器行为偏差?
主要发现
- FaceForensics++数据集严重失衡,61.7%的真实视频包含白人受试者,36.6%为白人女性。
- FaceForensics++中59.44%的虚假视频为'不规则'替换,即某人的面部被替换到不同种族或性别的对象上。
- Blended Images数据集中有65.45%为'不规则'替换,其中35.3%的亚裔女性面部被替换到白人女性面部上。
- 在这些数据集上训练的深度伪造检测器,对某些种族子群体的错误率最高上升10.7%,尤其是亚裔女性。
- 在Blended Images数据集上训练的检测器对亚裔女性受试者表现出系统性歧视,可能源于接触了大量不规则替换。
- 本研究发现,'不规则'面部替换在面部特征与虚假性之间制造了虚假相关性,从而损害了检测模型的公平性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。