Skip to main content
QUICK REVIEW

[论文解读] No computation without representation: Avoiding data and algorithm biases through diversity

Caitlin Kuhlman, Latifa Jackson|arXiv (Cornell University)|Feb 26, 2020
Ethics and Social Impacts of AI参考文献 61被引用 17
一句话总结

本文主张,若不从根本上多元化计算社区本身,就无法实现伦理AI;数据科学中代表性不足会加剧数据集和算法中的结构性偏见。通过有针对性地开展教育、指导与合作,特别是与少数族裔服务机构合作,将代表性不足的声音融入其中,可从源头嵌入公平性,减少算法偏见,推动建立更加公平的社会技术系统。

ABSTRACT

The emergence and growth of research on issues of ethics in AI, and in particular algorithmic fairness, has roots in an essential observation that structural inequalities in society are reflected in the data used to train predictive models and in the design of objective functions. While research aiming to mitigate these issues is inherently interdisciplinary, the design of unbiased algorithms and fair socio-technical systems are key desired outcomes which depend on practitioners from the fields of data science and computing. However, these computing fields broadly also suffer from the same under-representation issues that are found in the datasets we analyze. This disconnect affects the design of both the desired outcomes and metrics by which we measure success. If the ethical AI research community accepts this, we tacitly endorse the status quo and contradict the goals of non-discrimination and equity which work on algorithmic fairness, accountability, and transparency seeks to address. Therefore, we advocate in this work for diversifying computing as a core priority of the field and our efforts to achieve ethical AI practices. We draw connections between the lack of diversity within academic and professional computing fields and the type and breadth of the biases encountered in datasets, machine learning models, problem formulations, and interpretation of results. Examining the current fairness/ethics in AI literature, we highlight cases where this lack of diverse perspectives has been foundational to the inequity in treatment of underrepresented and protected group data. We also look to other professional communities, such as in law and health, where disparities have been reduced both in the educational diversity of trainees and among professional practices. We use these lessons to develop recommendations that provide concrete steps for the computing community to increase diversity.

研究动机与目标

  • 通过正视计算和数据科学社区中缺乏多样性这一根本原因,来解决算法偏见问题。
  • 强调AI研究中代表性不足导致在公平性、问责制和透明度方面的盲点。
  • 证明多样化的视角对于识别和缓解数据与模型设计中的结构性不平等至关重要。
  • 提出切实可行的策略,如指导、教育合作以及基于社区的研究协作,以构建更具包容性的AI研究生态系统。
  • 倡导一种自下而上、以社区为中心的多样性方法,使代表性不足群体能够成为伦理AI的领导者。

提出的方法

  • 分析现有算法公平性文献,识别因缺乏多样化视角而导致模型设计偏见的案例。
  • 将计算领域中代表性不足与数据集和算法中的系统性偏见进行类比。
  • 考察法律与医疗领域中成功的多样性举措,以借鉴计算领域的最佳实践。
  • 强调指导和亲和力工作坊(如霍华德大学的BPDM工作坊)在建立社区联系和提升技术能力方面的作用。
  • 提议与少数族裔服务机构(MSIs)开展教育合作,以提升AI研究人才管道中的代表性。
  • 倡导与社区领域专家建立研究合作,以确保模型反映现实世界的社会公平关切。

实验结果

研究问题

  • RQ1计算研究中某些族裔群体代表性不足,如何导致持续存在的算法偏见?
  • RQ2同质化的研究团队在哪些方面未能识别或解决数据与模型设计中的结构性不平等?
  • RQ3指导和社区建设项目在提升AI研究中的多样性与公平性方面发挥什么作用?
  • RQ4与少数族裔服务机构及社区专家的合作如何提升AI系统的公平性与相关性?
  • RQ5为实现AI与伦理计算中可持续的多样性,教育与研究文化需要哪些系统性变革?

主要发现

  • 计算社区中缺乏多样性直接导致了偏见数据集和算法的出现,因为同质化团队无法识别或考虑结构性不平等。
  • 多样化的视角对于识别和缓解算法系统中隐性与显性歧视至关重要,尤其是在未直接使用受保护属性的情况下。
  • 指导与亲和力工作坊(如霍华德大学的BPDM工作坊)成功促进了社区建设,提升了技术能力,并增加了代表性不足群体在数据科学中的参与度。
  • 与少数族裔服务机构的教育合作有助于培养更具代表性的AI与数据科学人才梯队。
  • 与代表性不足社区的领域专家开展基于社区的研究合作,可带来更加公平且符合语境的模型开发。
  • 若不主动努力多元化AI研究社区,仅靠算法调整来实现公平性的努力将始终不足且短暂。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。