[論文レビュー] Demographic-Reliant Algorithmic Fairness: Characterizing the Risks of Demographic Data Collection in the Pursuit of Fairness
この論文は、アルゴリズムの公平性を達成するために人種、性別、性的指向などの人口統計的データを収集することが、本質的に有益であるという仮定に挑戦し、こうしたデータ収集がシステム的抑圧を強化し、監視を可能にし、弱い立場のアイデンティティを誤って表現するリスクを伴うと主張する。本論文は、被害を軽減しながらAIシステムにおける公平性を推進するための責任ある代替策として、プライバシーを守るデータ収集と参加型ガバナンスモデルを提唱する。
Most proposed algorithmic fairness techniques require access to data on a "sensitive attribute" or "protected category" (such as race, ethnicity, gender, or sexuality) in order to make performance comparisons and standardizations across groups, however this data is largely unavailable in practice, hindering the widespread adoption of algorithmic fairness. Through this paper, we consider calls to collect more data on demographics to enable algorithmic fairness and challenge the notion that discrimination can be overcome with smart enough technical methods and sufficient data alone. We show how these techniques largely ignore broader questions of data governance and systemic oppression when categorizing individuals for the purpose of fairer algorithmic processing. In this work, we explore under what conditions demographic data should be collected and used to enable algorithmic fairness methods by characterizing a range of social risks to individuals and communities. For the risks to individuals we consider the unique privacy risks associated with the sharing of sensitive attributes likely to be the target of fairness analysis, the possible harms stemming from miscategorizing and misrepresenting individuals in the data collection process, and the use of sensitive data beyond data subjects' expectations. Looking more broadly, the risks to entire groups and communities include the expansion of surveillance infrastructure in the name of fairness, misrepresenting and mischaracterizing what it means to be part of a demographic group or to hold a certain identity, and ceding the ability to define for themselves what constitutes biased or unfair treatment. We argue that, by confronting these questions before and during the collection of demographic data, algorithmic fairness methods are more likely to actually mitigate harmful treatment disparities without reinforcing systems of oppression.
研究の動機と目的
- アルゴリズムの公平性のための人口統計的データ収集に関連する社会的・システム的リスクを検討すること。
- 人種、性別、性的指向に関するデータを増やすことで自動的にAIシステムがより公平になるという仮定に疑問を呈すること。
- 個人に対するリスク(例:プライバシー侵害、誤分類)と、コミュニティレベルでのリスク(例:監視インフラの拡大、抑圧的分類体系の強化)を特定すること。
- 参加型データガバナンスとプライバシー中心のデータ収集を、現在の実践の倫理的代替策として提唱すること。
- 弱い立場のコミュニティが公平性とデータガバナンスの定義に中心に立つ責任ある人口統計的データ使用の枠組みを提供すること。
提案手法
- グループのパフォーマンスを比較するために人口統計的データに依存する既存のアルゴリズム公平性技術を分析する。
- 個人に対するリスク(例:プライバシー侵害、誤表現、同意を超えたデータの悪用)を特定・分類する。
- 監視インfraストラクチャの拡大や抑圧的分類スキームの強化といったコミュニティレベルのリスクを検討する。
- 差分プライバシーとフェデレーテッドラーニングなどのプライバシーを守るデータ収集技術を評価し、感受性の高い属性の露出を低減する。
- 弱い立場のコミュニティが自らのアイデンティティを定義し、データ使用を制御できる参加型データガバナンスモデル(例:データ協同組合、信託)を提言する。
- コミュニティ主権に基づく倫理的データガバナンスを根拠づけるために、クリティカルレース理論と先住民のデータ主権を活用する。
実験結果
リサーチクエスチョン
- RQ1アルゴリズムの公平性のための人口統計的データ収集が、どのような条件下で倫理的に正当化可能か?
- RQ2感受性の高い属性(例:人種や性別)が収集されると、個人が直面するプライバシーリスクと誤分類の損害はどのように生じるか?
- RQ3人口統計的データ収集が、コミュニティレベルで監視を拡大し、システム的抑圧を強化するメカニズムは何か?
- RQ4参加型ガバナンスモデルは、弱い立場のグループが公平性を定義し、自身のデータを制御できるようにする仕組みは何か?
- RQ5中央集権的な人口統計的データに依存せずに、公平性を責任を持って実現するための技術的・制度的メカニズムは何か?
主な発見
- 公平性のための人口統計的データ収集は、差別のを必ずしも軽減せず、監視と誤表現を通じてシステム的不平等を固定化する可能性がある。
- 人種や性別といった感受性の高い属性が収集されると、特に元の同意を超えてデータが共有される場合、個人のプライバシーリスクが顕著に高まる。
- アイデンティティの誤分類や誤表現は、交差的アイデンティティを反映しないきめの粗いカテゴリーが適用される場合、排除やレッテル貼りを引き起こす可能性がある。
- 監視インフラは、データ対象者を越えて、一部のデータが他の人々のプロファイリングを可能にすることで、全体のコミュニティに影響を及ぼすことがある。
- 参加型ガバナンスモデル(例:データ協同組合、先住民のデータ主権フレームワーク)は、コミュニティが自らのアイデンティティを定義し、データ使用を制御できるようにする。
- 差分プライバシーやフェデレーテッドラーニングといったプライバシーを守る技術は、感受性の高いデータの露出を低減しつつ、公平性の評価を支援できる。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。