[论文解读] "What We Can't Measure, We Can't Understand": Challenges to Demographic Data Procurement in the Pursuit of Fairness
本文研究了从业者在获取算法公平性所需的人口统计学数据时面临的实际挑战,揭示了尽管此类数据至关重要,但法律、伦理和技术障碍常常导致无法获取。文章反对简单降低数据收集门槛的做法,转而倡导建立规范性框架和包容性实践,以在不直接依赖人口统计学数据收集的前提下,伦理地评估和缓解偏见。
As calls for fair and unbiased algorithmic systems increase, so too does the number of individuals working on algorithmic fairness in industry. However, these practitioners often do not have access to the demographic data they feel they need to detect bias in practice. Even with the growing variety of toolkits and strategies for working towards algorithmic fairness, they almost invariably require access to demographic attributes or proxies. We investigated this dilemma through semi-structured interviews with 38 practitioners and professionals either working in or adjacent to algorithmic fairness. Participants painted a complex picture of what demographic data availability and use look like on the ground, ranging from not having access to personal data of any kind to being legally required to collect and use demographic data for discrimination assessments. In many domains, demographic data collection raises a host of difficult questions, including how to balance privacy and fairness, how to define relevant social categories, how to ensure meaningful consent, and whether it is appropriate for private companies to infer someone's demographics. Our research suggests challenges that must be considered by businesses, regulators, researchers, and community groups in order to enable practitioners to address algorithmic bias in practice. Critically, we do not propose that the overall goal of future work should be to simply lower the barriers to collecting demographic data. Rather, our study surfaces a swath of normative questions about how, when, and whether this data should be procured, and, in cases where it is not, what should still be done to mitigate bias.
研究动机与目标
- 理解从业者在获取算法公平性所需人口统计学数据时面临的真实世界挑战。
- 考察影响工业界和研究领域中人口统计学数据使用之伦理、法律和技术限制。
- 评估替代性公平方法(如代理变量、推断或第三方审计)在多大程度上可有效替代直接的人口统计学数据收集。
- 探讨数据隐私法规与反歧视法律在实践中如何与公平性度量相互作用。
- 评估数据主体在公平性流程中的角色,以及基于同意的包容性数据收集模式的可行性。
提出的方法
- 对38名在算法公平性领域或相关领域工作的从业者及专业人员进行了半结构化访谈。
- 通过跨多样化领域的从业者经验分析,绘制出人口统计学数据可用性与使用情况的全谱图。
- 探讨了法律与监管框架,包括GDPR和美国反歧视法律,与数据获取的关系。
- 评估了差分隐私、密码学方法及第三方数据处理等隐私保护替代方案。
- 评估了联邦学习等去中心化方法在无需集中化人口统计学数据的前提下实现公平性分析的潜力。
- 考察了数据主体同意与社区参与在塑造公平性评估中伦理数据使用方面的角色。
实验结果
研究问题
- RQ1从业者在获取用于算法公平性评估的人口统计学数据时面临的主要障碍是什么?
- RQ2像GDPR这样的法律与监管框架在多大程度上影响了公平性工作中人口统计学数据的收集与使用?
- RQ3差分隐私或第三方数据处理等隐私保护技术在多大程度上可以替代直接的人口统计学数据收集?
- RQ4从业者在收集人口统计学数据时,如何应对关于同意、代表性及群体显著性等方面的伦理关切?
- RQ5在不依赖传统人口统计学数据收集的前提下,数据主体与社区利益相关者可在多大程度上参与并塑造公平性评估流程?
主要发现
- 由于法律限制、隐私担忧或组织政策,许多从业者即使在必要时也难以获取人口统计学数据,这些数据对偏见检测至关重要。
- GDPR等法律框架通常限制人口统计学数据的使用,但某些司法管辖区(如英国)已发布指导文件,允许在特定公平性审计条件下使用此类数据。
- 推断、代理变量或事后群体识别等替代方法可减少直接数据收集,但会引入误标示风险并降低公平性评估的准确性。
- 第三方数据收集及隐私保护技术(如差分隐私)虽转移了责任,但未必能解决关于数据显著性与群体代表性等核心伦理问题。
- 从业者对数据可靠性、同意机制以及通过有缺陷的人口统计分类可能加剧系统性偏见表示强烈担忧。
- 将数据主体与社区纳入公平性流程可提升数据的相关性与伦理问责性,尽管此类模式在实施上复杂且成本高昂。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。