Skip to main content
QUICK REVIEW

[論文レビュー] GBDF: Gender Balanced DeepFake Dataset Towards Fair DeepFake Detection

Aakash Varma Nadimpalli, Ajita Rattani|arXiv (Cornell University)|Jul 21, 2022
Face recognition and analysis被引用数 10
ひとこと要約

本稿では、公平性を評価するための手動による性別ラベルが付与された、ジェンダーに偏りのない深フェイクデータセットGBDFを紹介する。既存のデータセットは男性に偏っており、最先端の検出器は女性に対して性能が劣ることが判明したが、GBDFで学習させることで等誤差率(EER)における性能差を最大60%まで低減できることが示された。

ABSTRACT

Facial forgery by deepfakes has raised severe societal concerns. Several solutions have been proposed by the vision community to effectively combat the misinformation on the internet via automated deepfake detection systems. Recent studies have demonstrated that facial analysis-based deep learning models can discriminate based on protected attributes. For the commercial adoption and massive roll-out of the deepfake detection technology, it is vital to evaluate and understand the fairness (the absence of any prejudice or favoritism) of deepfake detectors across demographic variations such as gender and race. As the performance differential of deepfake detectors between demographic subgroups would impact millions of people of the deprived sub-group. This paper aims to evaluate the fairness of the deepfake detectors across males and females. However, existing deepfake datasets are not annotated with demographic labels to facilitate fairness analysis. To this aim, we manually annotated existing popular deepfake datasets with gender labels and evaluated the performance differential of current deepfake detectors across gender. Our analysis on the gender-labeled version of the datasets suggests (a) current deepfake datasets have skewed distribution across gender, and (b) commonly adopted deepfake detectors obtain unequal performance across gender with mostly males outperforming females. Finally, we contributed a gender-balanced and annotated deepfake dataset, GBDF, to mitigate the performance differential and to promote research and development towards fairness-aware deep fake detectors. The GBDF dataset is publicly available at: https://github.com/aakash4305/GBDF

研究の動機と目的

  • 性別にわたる深フェイク検出器の公平性の不均衡、特に女性に対する性能の低さを調査すること。
  • 公平性分析を妨げる既存の深フェイクデータセットにおける人種的属性の欠落を是正すること。
  • 公平性に配慮した深フェイク検出研究を支援するため、ジェンダーに偏りのない公開データセット(GBDF)を構築すること。
  • 実世界の深フェイク生成技術を用いて、データセットバイアスが検出器性能に与える影響を評価すること。
  • 可視化と指標を通じて、検出器が男性と女性で異なる顔面領域に注目していること、これが性能差を生じさせることを示すこと。

提案手法

  • 人間のレーティング担当者を用いて、FaceForensics++とCeleb-DFという2つの代表的な深フェイクデータセットの被験者を男性または女性に手動でラベル付けした。
  • 訓練データ内のジェンダー分布をバランスさせることで、男性と女性が等しく表現されるようにGBDFを構築した。
  • MesoInception-4、LipForensics、EfficientNet V2-Lなど、複数の最先端の深フェイク検出器を、元のデータセットおよびGBDFデータセットで訓練・評価した。
  • 勾配重み付きクラス活性化マッピング(Grad-CAM)を適用し、男性・女性の両方の性別において、ライブ画像とフェイク画像の分類に使用される注目領域を可視化・比較した。
  • 等誤差率(EER)を用いて公平性を定量化することで、男性および女性のサブグループ間での性能差を測定した。
  • データの不均衡が与える影響を分離するために、不均衡な(元の)データセットとバランスの取れた(GBDF)データセットの両方で検出器の性能を比較するアブレーションスタディを実施した。

実験結果

リサーチクエスチョン

  • RQ1既存の深フェイクデータセットにおいて、検出器の性能は男性と女性で顕著に異なるか?
  • RQ2現在の深フェイクデータセットにおけるジェンダーの不均衡は、検出モデルの性能差にどの程度寄与しているか?
  • RQ3GBDFのようなジェンダーに偏りのないデータセットで学習させることで、男性と女性の間の深フェイク検出における性能格差を是正できるか?
  • RQ4深フェイク検出器は、性別によって異なる顔面領域に依存して分類しているのか?その影響は公平性にどう現れるか?
  • RQ5深フェイク生成技術の選択(例:不規則な交換)は、性別による性能差にどのように影響するか?

主な発見

  • FaceForensics++ や Celeb-DF といった既存の深フェイクデータセットは、実画像およびフェイク画像の両方において顕著なジェンダーの不均衡を示しており、女性が過小代表されている。
  • 特に MesoInception-4 は、不均衡なデータセット上で男性と女性の間で等誤差率(EER)に最大 0.034 の性能差を示した。
  • ジェンダー分布がバランスされた GBDF データセットでは、GBDF テストセットで EER の差が 0.031 から 0.012 にまで低下した。
  • Grad-CAM の可視化から、検出器が女性に対しては頬、男性に対しては目の周辺領域に注目していることが判明し、性別に依存した特徴の利用が示された。
  • 口元の動きに注目し、クロップされた領域を用いる LipForensics などのモデルは、性別間で最も小さな性能差を示しており、モダリティの選択がバイアスを軽減できる可能性を示唆している。
  • FaceForensics++ において女性のサンプル数が男性より多い場合でも性能格差が継続していたことから、不均衡が唯一の要因ではないことが判明。モデルバイアスやデータ品質の要因も関与している。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。