Skip to main content
QUICK REVIEW

[论文解读] Towards Accountability for Machine Learning Datasets: Practices from Software Engineering and Infrastructure

Ben Hutchinson, Andrew Smart|arXiv (Cornell University)|Oct 23, 2020
Ethics and Social Impacts of AI参考文献 129被引用 46
一句话总结

本文认为 ML 数据集是基础设施级的产物,并提出一个以生命周期为基础、严格的文档框架,借鉴软件工程,以确保透明度、问责制和负责任的数据集开发。

ABSTRACT

Rising concern for the societal implications of artificial intelligence systems has inspired demands for greater transparency and accountability. However the datasets which empower machine learning are often used, shared and re-used with little visibility into the processes of deliberation which led to their creation. Which stakeholder groups had their perspectives included when the dataset was conceived? Which domain experts were consulted regarding how to model subgroups and other phenomena? How were questions of representational biases measured and addressed? Who labeled the data? In this paper, we introduce a rigorous framework for dataset development transparency which supports decision-making and accountability. The framework uses the cyclical, infrastructural and engineering nature of dataset development to draw on best practices from the software development lifecycle. Each stage of the data development lifecycle yields a set of documents that facilitate improved communication and decision-making, as well as drawing attention the value and necessity of careful data work. The proposed framework is intended to contribute to closing the accountability gap in artificial intelligence systems, by making visible the often overlooked work that goes into dataset creation.

研究动机与目标

  • 主张 ML 数据集作为需要可视性和问责性的技术基础设施来运作。
  • 倡导在数据集开发中采用软件工程的生命周期实践。
  • 提出一个结构化的文档模型,包含特定的产出物类型,以实现审计和评估。
  • 突出数据集工作中的政治与工程维度及其非线性生命周期。

提出的方法

  • 将数据集框架为基础设施和工程产物,以证明对问责的需求。
  • 将数据集开发阶段映射到类似软件的生命周期(需求、设计、实现、测试、维护)。
  • 在每个阶段引入文档类型,以促进可追溯性和问责制(Requirements Analysis Documents, Dataset Design Documents, Implementation Diaries, Testing Reports, Maintenance Plans)。
  • 提出治理概念,如审计、多元监督和事后分析,以弥合问责差距。

实验结果

研究问题

  • RQ1在数据集开发的事前、进行中和事后,应该记录哪些信息以实现有意义的问责?
  • RQ2如何改编软件工程实践以提升 ML 数据集的可见性、所有权与可审计性?
  • RQ3在数据集开发生命周期中,需要哪些关键的文档产出物及所有权角色?
  • RQ4将数据集视为基础设施的概念如何影响 ML 的问责性与治理?

主要发现

  • 数据集最好被视为能够使 ML 系统运作的基础设施,因此需要经过深思熟虑、不过于匆忙的开发与文档化。
  • 一个非线性、迭代的数据集开发生命周期,具备明确的所有权与文档化,可减少问责差距。
  • 在每个生命周期阶段(需求、设计、实现、测试、维护)有一套结构化的文档,有助于可追溯性与问责。
  • 审计、评审和持续维护计划对于解决数据陈旧、错误和情境变化至关重要。
  • 文档应明确反映假设、权衡取舍和利益相关者的讨论,以对抗偏见和意外伤害。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。