Skip to main content
QUICK REVIEW

[论文解读] Healthsheet: Development of a Transparency Artifact for Health Datasets

Negar Rostamzadeh, Diana Mincu|arXiv (Cornell University)|Feb 26, 2022
Mobile Health and mHealth Applications被引用 8
一句话总结

本文提出了 Healthsheet,作为数据集文档框架(datasheets for datasets)在医疗保健领域的专用适配版本,旨在提升机器学习驱动的医疗研究中的透明度与问责性。通过针对电子健康记录(EHR)、临床试验和数字健康数据集的访谈与案例研究,作者证明 Healthsheet 能够提升偏差检测能力,支持伦理数据集评估,并促进社区驱动的文档编制,从而推动医疗领域中公平的机器学习应用。

ABSTRACT

Machine learning (ML) approaches have demonstrated promising results in a wide range of healthcare applications. Data plays a crucial role in developing ML-based healthcare systems that directly affect people's lives. Many of the ethical issues surrounding the use of ML in healthcare stem from structural inequalities underlying the way we collect, use, and handle data. Developing guidelines to improve documentation practices regarding the creation, use, and maintenance of ML healthcare datasets is therefore of critical importance. In this work, we introduce Healthsheet, a contextualized adaptation of the original datasheet questionnaire ~\cite{gebru2018datasheets} for health-specific applications. Through a series of semi-structured interviews, we adapt the datasheets for healthcare data documentation. As part of the Healthsheet development process and to understand the obstacles researchers face in creating datasheets, we worked with three publicly-available healthcare datasets as our case studies, each with different types of structured data: Electronic health Records (EHR), clinical trial study data, and smartphone-based performance outcome measures. Our findings from the interviewee study and case studies show 1) that datasheets should be contextualized for healthcare, 2) that despite incentives to adopt accountability practices such as datasheets, there is a lack of consistency in the broader use of these practices 3) how the ML for health community views datasheets and particularly extit{Healthsheets} as diagnostic tool to surface the limitations and strength of datasets and 4) the relative importance of different fields in the datasheet to healthcare concerns.

研究动机与目标

  • 解决机器学习中使用的医疗数据集缺乏标准化、伦理化文档的问题。
  • 识别当前医疗领域机器学习社区中数据文档实践存在的障碍与不一致之处。
  • 开发一种针对医疗数据独特需求与挑战的上下文化数据集文档框架。
  • 通过结构化透明度,使数据使用者能够更好地评估数据质量、偏差与代表性。
  • 促进社区对数据集文档的采纳与持续演进,作为医疗领域中伦理化与可问责机器学习的工具。

提出的方法

  • 改编原始的数据集文档框架(Gebru et al., 2018)以创建 Healthsheet,一种面向医疗保健的透明度工具。
  • 对研究人员和数据管理人进行半结构化访谈,以识别关键的文档需求与挑战。
  • 将 Healthsheet 应用于三个公开可用的医疗数据集:MIMIC-III(电子健康记录)、临床试验数据以及基于智能手机的性能测量数据。
  • 将数据集核心组件(如数据收集、版本控制和创建者人口统计)映射至医疗保健特定的关注点。
  • 通过迭代式案例研究与利益相关者反馈,评估 Healthsheet 的可行性与实用性。
  • 提出一种社区驱动的模型,用于在新数据集版本和应用场景中持续维护与演进 Healthsheet。

实验结果

研究问题

  • RQ1如何使数据集文档适应医疗数据特有的伦理与结构挑战?
  • RQ2研究人员在创建和维护医疗数据集全面文档时面临的主要障碍是什么?
  • RQ3数据集使用者如何看待透明度工具(如 Healthsheet)在评估数据质量与偏差方面的价值?
  • RQ4在医疗相关机器学习应用中,数据集的哪些组成部分对伦理评估最为关键?
  • RQ5Healthsheet 在多大程度上可作为诊断工具,用于识别现实医疗数据集中数据局限性与优势?

主要发现

  • Healthsheet 有效将通用数据集文档原则情境化应用于医疗领域,回应了数据来源、偏差与临床相关性等特定领域关切。
  • 尽管普遍认可透明度的价值,但当前医疗数据集的文档实践仍存在显著不一致。
  • 研究人员将 Healthsheet 视为一种诊断工具,可增强对数据集局限性与优势的理解,尤其在识别偏差与数据代表性方面。
  • 完成 Healthsheet 的过程揭示了现有元数据中的重大缺失,特别是关于数据获取机制与版本控制方案的信息。
  • 社区验证与 Healthsheet 条目持续维护至关重要,因为只有数据集所有者才能准确报告数据意图与来源。
  • 缺乏数据集版本的标准化命名规范,以及缺乏集中化的文档存储库,严重阻碍了医疗机器学习研究中的透明度与可复现性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。