Skip to main content
QUICK REVIEW

[论文解读] Biases in Data Science Lifecycle

Dinh-An Ho, Oya Beyan|arXiv (Cornell University)|Sep 10, 2020
Ethics and Social Impacts of AI参考文献 54被引用 9
一句话总结

本文识别并分类了数据科学生命周期各个阶段中的偏见,为数据科学家提供了一个实用的、分阶段的框架,以识别和减轻伦理风险。通过将偏见来源映射到数据收集、预处理、建模和部署等阶段,本研究为减少现实应用中的意外后果提供了可操作的指导。

ABSTRACT

In recent years, data science has become an indispensable part of our society. Over time, we have become reliant on this technology because of its opportunity to gain value and new insights from data in any field - business, socializing, research and society. At the same time, it raises questions about how justified we are in placing our trust in these technologies. There is a risk that such powers may lead to biased, inappropriate or unintended actions. Therefore, ethical considerations which might occur as the result of data science practices should be carefully considered and these potential problems should be identified during the data science lifecycle and mitigated if possible. However, a typical data scientist has not enough knowledge for identifying these challenges and it is not always possible to include an ethics expert during data science production. The aim of this study is to provide a practical guideline to data scientists and increase their awareness. In this work, we reviewed different sources of biases and grouped them under different stages of the data science lifecycle. The work is still under progress. The aim of early publishing is to collect community feedback and improve the curated knowledge base for bias types and solutions.

研究动机与目标

  • 提高数据科学家对数据科学工作流中由偏见引发的伦理风险的认识。
  • 系统性地识别数据科学生命周期各阶段的偏见来源。
  • 为数据科学家提供一个实用且易于访问的框架,使其能够主动检测并解决偏见,而无需具备伦理专业知识。
  • 通过早期发布一个精心整理的偏见类型与缓解策略知识库,促进社区反馈。
  • 支持在各行业和领域内发展更加负责任和公平的数据科学实践。

提出的方法

  • 对现有的数据科学伦理与偏见相关文献进行了全面综述。
  • 将识别出的偏见来源映射到数据科学生命周期的五个核心阶段:问题定义、数据收集、数据预处理、建模和部署。
  • 根据偏见的起源和在各生命周期阶段的影响,将其归类为主题性类别。
  • 采用结构化、分阶段的方法,组织偏见类型及其相应的缓解策略。
  • 强调针对非专业数据科学家量身定制的实用、可操作的指导。
  • 以预印本形式发布研究成果,以征求社区反馈,并通过迭代方式改进知识库。

实验结果

研究问题

  • RQ1在数据科学生命周期的不同阶段,哪些类型的偏见最有可能出现?
  • RQ2没有正式伦理培训的数据科学家如何识别并解决其工作流中的偏见?
  • RQ3可以采用何种系统性框架,将偏见来源映射到数据科学生命周期中,以实现实际的缓解?
  • RQ4社区反馈如何改善数据科学中偏见类型与解决方案的整理工作?
  • RQ5在标准实践中常被忽视的数据科学关键伦理风险是什么?

主要发现

  • 偏见贯穿于数据科学生命周期的所有阶段,从问题定义到部署。
  • 数据收集和预处理阶段尤其容易出现抽样偏见和测量偏见。
  • 建模阶段常因假设和数据表示问题而引入算法偏见和特征选择偏见。
  • 部署和监控阶段可能通过反馈回路和可解释性不足而放大偏见。
  • 本研究在生命周期中识别出12种不同的偏见类型,每种均附有示例和缓解策略。
  • 早期社区反馈已促成偏见分类体系的优化,并提升了框架的实际可用性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。