[论文解读] Reproducibility of the Methods in Medical Imaging with Deep Learning
本文通过分析2018至2022年MIDL会议的所有投稿,评估了医学影像深度学习方法的可重现性,发现仅有22%的代码仓库被认为可重复。论文提出了一套可重现性检查清单,并建议会议层面制定指南,以提升代码仓库质量、代码透明度和数据引用,从而增强医学人工智能研究的方法严谨性和影响力。
Concerns about the reproducibility of deep learning research are more prominent than ever, with no clear solution in sight. The relevance of machine learning research can only be improved if we also employ empirical rigor that incorporates reproducibility guidelines, especially so in the medical imaging field. The Medical Imaging with Deep Learning (MIDL) conference has made advancements in this direction by advocating open access, and recently also recommending authors to make their code public - both aspects being adopted by the majority of the conference submissions. This helps the reproducibility of the methods, however, there is currently little or no support for further evaluation of these supplementary material, making them vulnerable to poor quality, which affects the impact of the entire submission. We have evaluated all accepted full paper submissions to MIDL between 2018 and 2022 using established, but slightly adjusted guidelines on reproducibility and the quality of the public repositories. The evaluations show that publishing repositories and using public datasets are becoming more popular, which helps traceability, but the quality of the repositories has not improved over the years, leaving room for improvement in every aspect of designing repositories. Merely 22% of all submissions contain a repository that were deemed repeatable using our evaluations. From the commonly encountered issues during the evaluations, we propose a set of guidelines for machine learning-related research for medical imaging applications, adjusted specifically for future submissions to MIDL.
研究动机与目标
- 解决深度学习研究中日益增长的可重现性担忧,特别是在方法论严谨性至关重要的医学影像领域。
- 调查2018至2022年MIDL会议投稿中公开共享代码仓库的质量与可重现性。
- 识别在仓库设计、数据引用和代码可用性方面阻碍可重现性的常见问题,尽管已有开放科学政策。
- 为作者和会议组织者提出可操作的指南,以提升补充研究材料的透明度、质量与长期可用性。
- 支持将在同行评审过程中纳入可重现性评估,以确保发表前的方法论稳健性。
提出的方法
- 使用针对医学影像和深度学习优化的可重现性检查清单,评估了2018至2022年MIDL会议所有被接受的完整论文投稿。
- 根据代码完整性、文档质量、环境配置、结果可重现性以及公共数据集的使用情况,对代码仓库进行评估。
- 审查公开数据集是否正确引用、是否提供可访问链接,以及是否与研究声明一致。
- 识别出常见问题,如链接失效、依赖项缺失、缺少训练/推理脚本,以及缺乏版本控制。
- 提出一个结构化的可重现性检查清单(附录A),并建议在OpenReview投稿表单中增加可选字段,以标准化仓库质量。
- 建议MIDL采用正式的认可系统,对满足所有检查清单标准的投稿进行认证,类似于ECML-PKDD的“可重现”标志。
实验结果
研究问题
- RQ1MIDL投稿所附代码仓库在多大程度上真正具备可重复性和可重现性?
- RQ2医学影像深度学习研究中,公开共享代码仓库在设计与维护方面最常见的缺陷是什么?
- RQ3使用公共数据集在多大程度上影响可重现性?这些数据集是否在投稿中得到一致引用和链接?
- RQ4结构化指南与会议层面的政策是否能显著提升补充材料的质量与透明度?
- RQ5同行评审在确保主论文内容之外的可重现性标准方面,应发挥何种作用?
主要发现
- 仅有22%的提交仓库被认为可重复,表明尽管代码共享日益普及,可重现性方面仍存在显著差距。
- 尽管公共数据集使用增加,但许多投稿未能正确引用或链接数据集,降低了可追溯性与透明度。
- 大量仓库为空或包含失效链接,其中9个空仓库中有4个在会议结束后一个月内被修改。
- 许多仓库缺少关键组件,如训练脚本、推理代码、环境配置以及清晰的文档。
- 研究表明,仅共享代码并不能保证可重现性;仓库的质量与可维护性是关键因素。
- 作者建议将在同行评审中纳入可重现性评估,并引入正式认可系统,以表彰高质量投稿。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。