Skip to main content
QUICK REVIEW

[论文解读] Awareness in Practice: Tensions in Access to Sensitive Attribute Data for Antidiscrimination

Miranda Bogen, Aaron Rieke|arXiv (Cornell University)|Dec 12, 2019
Ethics and Social Impacts of AI参考文献 37被引用 13
一句话总结

本文研究了美国信贷、就业和医疗保健领域在现实世界中对敏感属性数据(如种族和性别)访问的不一致性,揭示了法律规范、企业实践和机构规范如何构成实施公平意识机器学习的重大障碍。文章认为,若无法可靠获取此类数据,偏差检测与缓解技术便无法在研究环境之外实际部署。

ABSTRACT

Organizations cannot address demographic disparities that they cannot see. Recent research on machine learning and fairness has emphasized that awareness of sensitive attributes, such as race and sex, is critical to the development of interventions. However, on the ground, the existence of these data cannot be taken for granted. This paper uses the domains of employment, credit, and healthcare in the United States to surface conditions that have shaped the availability of sensitive attribute data. For each domain, we describe how and when private companies collect or infer sensitive attribute data for antidiscrimination purposes. An inconsistent story emerges: Some companies are required by law to collect sensitive attribute data, while others are prohibited from doing so. Still others, in the absence of legal mandates, have determined that collection and imputation of these data are appropriate to address disparities. This story has important implications for fairness research and its future applications. If companies that mediate access to life opportunities are unable or hesitant to collect or infer sensitive attribute data, then proposed techniques to detect and mitigate bias in machine learning models might never be implemented outside the lab. We conclude that today's legal requirements and corporate practices, while highly inconsistent across domains, offer lessons for how to approach the collection and inference of sensitive data in appropriate circumstances. We urge stakeholders, including machine learning practitioners, to actively help chart a path forward that takes both policy goals and technical needs into account.

研究动机与目标

  • 调查为何对敏感属性数据的访问在关键美国领域中不一致,而这些数据对检测和缓解机器学习中的偏差至关重要。
  • 分析法律要求、企业政策和机构规范如何塑造信贷、就业和医疗保健领域中敏感属性的收集与推断。
  • 强调若未收集或推断敏感数据,公平意识机器学习技术将无法在实践中实施的风险。
  • 呼吁研究人员、政策制定者和从业者开展跨学科合作,为公平干预中的负责任数据使用建立明确原则。
  • 探讨推断敏感属性的伦理与技术挑战,并提出其负责任处理的保障措施。

提出的方法

  • 对美国三个受监管领域(信贷、就业和医疗保健)中敏感属性数据收集的法律与制度背景进行比较分析。
  • 绘制各领域中收集敏感属性的现有法律要求与禁止规定,识别出不同的监管框架。
  • 基于合规性、道德义务或风险缓解需求,探讨各领域中私营企业如何收集或推断敏感属性以实现反歧视目的。
  • 评估数据最小化和隐私法律对公平研究的影响,特别是在机器学习模型审计的背景下。
  • 提出技术保障措施,如安全多方计算、私有集合交集和同态加密,以实现对敏感数据的隐私保护访问。
  • 倡导制定明确的政策指导,说明在何种情况下以及如何收集或推断敏感属性,包括对推断数据进行标记并单独存储。

实验结果

研究问题

  • RQ1在美国信贷、就业和医疗保健领域,私营企业基于反歧视目的在何种法律与制度条件下收集或推断敏感属性数据?
  • RQ2为何一些公司收集敏感属性数据,而另一些公司尽管目标相似(减少歧视)却受到禁止?
  • RQ3隐私法律与数据最小化原则如何与公平意识机器学习中对受保护群体认知的需求产生冲突?
  • RQ4哪些技术和伦理保障措施可确保在机器学习系统中负责任地收集和使用推断的敏感属性数据?
  • RQ5机器学习研究人员和从业者在塑造未来敏感属性数据访问政策方面应发挥何种作用?

主要发现

  • 在美国信贷领域,部分贷款机构依法必须收集敏感属性数据,而其他机构则被禁止,导致数据收集格局支离破碎。
  • 大型雇主通常将敏感属性数据收集作为标准人力资源实践的一部分,反映出长期遵守民权法律的惯例。
  • 在医疗保健领域,企业不仅出于法律合规,还出于减少种族和性别间显著健康结果差异的道德义务而推动数据收集。
  • 技术平台和供应商——尽管在决定人生机会方面发挥中介作用——却缺乏关于收集或推断敏感属性的明确指导,导致数据收集不一致或完全缺失。
  • 目前尚无一致且广泛接受的框架来确定何时或如何收集或推断敏感属性数据,尤其针对非传统受监管实体。
  • 隐私保护技术如同态加密和私有集合交集虽具前景,但尚不成熟,尚未在公平性测试中广泛部署。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。