Skip to main content
QUICK REVIEW

[论文解读] Building Better Datasets: Seven Recommendations for Responsible Design from Dataset Creators

Will Orr, Kate Crawford|arXiv (Cornell University)|Aug 30, 2024
Big Data and Business Intelligence被引用 7
一句话总结

本文从18位数据集创建者那里获得见解,并提出七条切实可行的建议,以改进负责任的数据集设计,重点关注数据质量、多样性、文档记录和治理。

ABSTRACT

The increasing demand for high-quality datasets in machine learning has raised concerns about the ethical and responsible creation of these datasets. Dataset creators play a crucial role in developing responsible practices, yet their perspectives and expertise have not yet been highlighted in the current literature. In this paper, we bridge this gap by presenting insights from a qualitative study that included interviewing 18 leading dataset creators about the current state of the field. We shed light on the challenges and considerations faced by dataset creators, and our findings underscore the potential for deeper collaboration, knowledge sharing, and collective development. Through a close analysis of their perspectives, we share seven central recommendations for improving responsible dataset creation, including issues such as data quality, documentation, privacy and consent, and how to mitigate potential harms from unintended use cases. By fostering critical reflection and sharing the experiences of dataset creators, we aim to promote responsible dataset creation practices and develop a nuanced understanding of this crucial but often undervalued aspect of machine learning research.

研究动机与目标

  • 突出数据集创建者在负责任ML实践中的作用和观点。
  • 识别跨越多样数据集与机构情境的共同挑战。
  • 提出可执行的建议,以提高数据质量、多样性、同意与使用限制。
  • 鼓励机器学习社区内数据工作之间的协作与专业化。

提出的方法

  • 于2022年7月至9月对18位数据集创建者进行了开放式定性访谈。
  • 招募了47名潜在参与者;18人参与(38%响应率)。
  • 使用半结构化访谈探讨数据集的起源、使用、维护与过时情况。
  • 对访谈进行了转录,并进行迭代主题编码以提取建议。
  • 确保参与者的匿名性或身份选择;研究获得伦理委员会批准。

实验结果

研究问题

  • RQ1跨领域的数据集创建者面临的挑战与最佳实践是什么?
  • RQ2创建者提出哪些具体建议来改进负责任的数据集创建?
  • RQ3数据集社区如何实现更好的协作与专业化?
  • RQ4文档记录、多样性、验证和同意在负责任的数据集设计中扮演什么角色?
  • RQ5数据集应如何被传达、使用,以及随时间可能的淘汰或更新?

主要发现

  • 访谈中出现了七条关于负责任的数据集创建的核心建议。
  • 多样性和全面审计对于减轻偏见和非预期伤害至关重要。
  • 高数据质量需要谨慎的验证、人工检查与整合,同时要意识到取舍。
  • 及早、迭代式的发展并从错误中学习是必不可少的。
  • 开放文档和清晰的局限性沟通支持可重复性与再利用。
  • 数据集应以用户为中心,明确界定预期用途并考虑意外用途。
  • 伦理考量、同意、隐私、许可、署名和淘汰等问题仍然是需要持续关注的开放挑战。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。