[论文解读] Demographic-Reliant Algorithmic Fairness: Characterizing the Risks of Demographic Data Collection in the Pursuit of Fairness
本文挑战了收集人口统计数据可天然促进算法公平性的假设,指出此类数据收集可能加剧系统性压迫,助长监控,并错误代表边缘化群体的身份。文章提出应采用隐私保护的数据收集方式和参与式治理模式,作为负责任的替代方案,以减轻危害并推动人工智能系统中的公平性。
Most proposed algorithmic fairness techniques require access to data on a "sensitive attribute" or "protected category" (such as race, ethnicity, gender, or sexuality) in order to make performance comparisons and standardizations across groups, however this data is largely unavailable in practice, hindering the widespread adoption of algorithmic fairness. Through this paper, we consider calls to collect more data on demographics to enable algorithmic fairness and challenge the notion that discrimination can be overcome with smart enough technical methods and sufficient data alone. We show how these techniques largely ignore broader questions of data governance and systemic oppression when categorizing individuals for the purpose of fairer algorithmic processing. In this work, we explore under what conditions demographic data should be collected and used to enable algorithmic fairness methods by characterizing a range of social risks to individuals and communities. For the risks to individuals we consider the unique privacy risks associated with the sharing of sensitive attributes likely to be the target of fairness analysis, the possible harms stemming from miscategorizing and misrepresenting individuals in the data collection process, and the use of sensitive data beyond data subjects' expectations. Looking more broadly, the risks to entire groups and communities include the expansion of surveillance infrastructure in the name of fairness, misrepresenting and mischaracterizing what it means to be part of a demographic group or to hold a certain identity, and ceding the ability to define for themselves what constitutes biased or unfair treatment. We argue that, by confronting these questions before and during the collection of demographic data, algorithmic fairness methods are more likely to actually mitigate harmful treatment disparities without reinforcing systems of oppression.
研究动机与目标
- 探讨为实现算法公平性而收集人口统计数据所伴随的社会与系统性风险。
- 挑战关于种族、性别和性取向等数据越多,人工智能系统就越公平的假设。
- 识别对个人的风险,如隐私泄露和误分类,以及对社区的风险,包括监控范围扩大和群体身份误代表。
- 倡导采用参与式数据治理和以隐私为核心的数据显示收集方式,作为当前做法的伦理替代方案。
- 提出一个以边缘化社区为中心、由其定义公平性与数据治理的负责任人口统计数据使用框架。
提出的方法
- 分析依赖人口统计数据来比较群体表现的现有算法公平性技术。
- 识别并分类对个人的风险,包括隐私侵犯、身份误代表以及超出原始同意范围的数据滥用。
- 探讨社区层面的风险,如监控基础设施扩张和压迫性分类体系的强化。
- 评估隐私保护的数据收集技术,如差分隐私和联邦学习,以减少敏感属性的暴露。
- 提出参与式数据治理模式——如数据合作社和数据信托——使边缘化社区能够自主定义自身身份并控制数据使用。
- 借鉴批判性种族理论和原住民数据主权理论,将伦理化数据治理建立在社区自主决定的基础之上。
实验结果
研究问题
- RQ1在何种条件下,为算法公平性而收集人口统计数据在伦理上是可接受的?
- RQ2当收集种族或性别等敏感属性时,隐私风险和误分类如何对个人造成伤害?
- RQ3人口统计数据的收集在哪些方面会扩大监控并从社区层面强化系统性压迫?
- RQ4参与式治理模式如何赋能边缘化群体自主定义公平性并掌控其数据?
- RQ5哪些技术和制度机制可以在不依赖集中化人口统计数据的前提下,负责任地实现公平性?
主要发现
- 为公平性而收集人口统计数据并不必然减少歧视,反而可能通过监控和身份误代表加剧系统性不公。
- 当种族或性别等敏感属性被收集时,个人面临更高的隐私风险,尤其是当数据在原始同意范围外被共享时。
- 身份误分类和误代表可能导致排斥或污名化,特别是当僵化的分类无法反映交叉性身份时。
- 监控基础设施可能从数据主体扩展至整个社区,尤其当部分群体的数据被用于对其他群体进行群体画像时。
- 参与式治理模式,如数据合作社和原住民数据主权框架,使社区能够自主定义自身身份并控制数据使用。
- 隐私保护技术如差分隐私和联邦学习可在减少敏感数据暴露的同时,仍支持公平性评估。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。