[論文レビュー] Auditing ImageNet: Towards a Model-driven Framework for Annotating Demographic Attributes of Large-Scale Image Datasets
この論文は、大規模な画像データセットにおける顔の年齢および性別の自動アノテーションを目的としたモデル駆動フレームワークを提案し、2012年 ImageNet ILSVRCサブセットおよび「person」シンセットに適用している。顔の性別のうち41.62%が女性、60歳以上はたったの1.71%であり、15〜29歳の男性が最も大きなグループであるなど顕著なデモグラフィックな不均衡が明らかになった。また、モデル自体が特に肌の色が濃い女性に対してバイアスを生じさせることも指摘された。
The ImageNet dataset ushered in a flood of academic and industry interest in deep learning for computer vision applications. Despite its significant impact, there has not been a comprehensive investigation into the demographic attributes of images contained within the dataset. Such a study could lead to new insights on inherent biases within ImageNet, particularly important given it is frequently used to pretrain models for a wide variety of computer vision tasks. In this work, we introduce a model-driven framework for the automatic annotation of apparent age and gender attributes in large-scale image datasets. Using this framework, we conduct the first demographic audit of the 2012 ImageNet Large Scale Visual Recognition Challenge (ILSVRC) subset of ImageNet and the "person" hierarchical category of ImageNet. We find that 41.62% of faces in ILSVRC appear as female, 1.71% appear as individuals above the age of 60, and males aged 15 to 29 account for the largest subgroup with 27.11%. We note that the presented model-driven framework is not fair for all intersectional groups, so annotation are subject to bias. We present this work as the starting point for future development of unbiased annotation models and for the study of downstream effects of imbalances in the demographics of ImageNet. Code and annotations are available at: http://bit.ly/ImageNetDemoAudit
研究の動機と目的
- コンピュータビジョン分野の基盤的データセットである ImageNet における包括的でないデモグラフィック分析の欠如に応えること。
- 特に性別および年齢の表現に関する、ImageNet の構成に内在するバイアスを調査すること。
- 大規模データセットにおける顔の年齢および性別の自動アノテーションのためのモデル駆動パイプラインを開発・適用すること。
- 今後の公平で包括的なデモグラフィックアノテーションモデル開発のためのベースラインを確立すること。
- ImageNet のデモグラフィック不均衡がトランスファー学習の過程でどのように拡大するかを、下流の分析を可能にすること。
提案手法
- 信頼度閾値 0.9 を設定した深層学習ベースの顔検出モデルを用い、画像内の顔を同定する。
- 性別アノテーションに DEX モデルを採用し、0 から 1 の連続値を出力し、0.5 を閾値として二値分類を実施する。
- 顔の年齢推定のための別個の深層学習モデルを適用し、APP A-REAL テストセットで評価する。
- ベンチマークデータセットを用いたモデル検証:顔検出には FDDB、年齢推定には APPA-REAL を使用する。
- パイLOT議会ベンチマーク(PPB)を用いたモデルの公平性評価により、交差集団間でのパフォーマンスの差が明らかになった。
- パイプラインを ILSVRC 2012 サブセット(128万枚の画像)および「person」階層的シンセット(118万枚の画像)に適用し、デモグラフィックアノテーションを生成した。
実験結果
リサーチクエスチョン
- RQ1ImageNet ILSVRC 2012 サブセットおよび「person」シンセットにおける顔のデモグラフィック構成(顔の年齢および性別)はどのようになっているか?
- RQ2既存の年齢および性別アノテーションの自動モデルは、特に交差集団カテゴリにおいて、さまざまなデモグラフィックサブグループでどの程度の性能を示すか?
- RQ3現在のモデル駆動型アノテーションフレームワークは、デモグラフィックラベル付けにおいてどの程度バイアスを導入または拡大するか?
- RQ4ImageNet の「person」と ILSVRC サブセットにおいて、最も過小および過剰に代表されているデモグラフィックグループは何か?
- RQ5モデル駆動フレームワークを、大規模なビジョンデータセットにおける公平性の監査および向上の基盤としてどのように活用できるか?
主な発見
- ILSVRC 2012 サブセットでは、検出可能な顔の41.62%が女性とアノテートされており、15〜29歳の男性が27.11%を占め、最も大きなグループである。
- ILSVRC で60歳以上の顔はたったの1.71%であり、高齢者の顕著な過小代表が示された。
- 性別アノテーションモデルはテストセットで平均適合率99.74%を達成したが、肌の色が濃い女性ではパフォーマンスが著しく低く(PPBで69.00%の正答率)、問題が明らかになった。
- 年齢推定モデルの平均平均誤差は全グループで5.22年であり、60歳以上では誤差が8.40年に上昇した。
- 「person」サブセットの上位シンセットでは極端な性別不均衡が見られ、特に「spree」、「bombshell」、「pitchman」は100%男性とアノテートされた。
- モデル駆動フレームワークにより、既存の自動アノテーションツールが交差集団間で公平でないことが判明し、特に肌の色が濃い女性に対しては今後の改善が不可欠であることが示された。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。