[论文解读] Reproducibility in Machine Learning-Driven Research
这篇小型调查研究探讨了机器学习(ML)驱动研究中的可复现性挑战,识别出关键障碍,如未公开的代码、数据泄露以及对训练条件的敏感性。研究评估了代码共享、模型信息表、检查清单、意识宣传和期刊政策等驱动因素——发现大型语言模型(如GPT-4)可实现86.6%的准确率来验证检查清单,且结构化实践显著提升了可复现性的可行性。
Research is facing a reproducibility crisis, in which the results and findings of many studies are difficult or even impossible to reproduce. This is also the case in machine learning (ML) and artificial intelligence (AI) research. Often, this is the case due to unpublished data and/or source-code, and due to sensitivity to ML training conditions. Although different solutions to address this issue are discussed in the research community such as using ML platforms, the level of reproducibility in ML-driven research is not increasing substantially. Therefore, in this mini survey, we review the literature on reproducibility in ML-driven research with three main aims: (i) reflect on the current situation of ML reproducibility in various research fields, (ii) identify reproducibility issues and barriers that exist in these research fields applying ML, and (iii) identify potential drivers such as tools, practices, and interventions that support ML reproducibility. With this, we hope to contribute to decisions on the viability of different solutions for supporting ML reproducibility.
研究动机与目标
- 评估不同研究领域中机器学习可复现性的当前状态。
- 识别系统性障碍,如未公开代码、数据泄露以及超参数敏感性,这些因素阻碍了机器学习研究的可复现性。
- 评估现有驱动因素,包括检查清单、模型信息表、意识倡议和期刊政策,以支持可复现性。
- 研究新兴工具(如大型语言模型)在高准确率下验证可复现性检查清单中的作用。
- 为评估不同可复现性解决方案在机器学习研究中的可行性提供基础。
提出的方法
- 开展系统性文献综述,聚焦于机器学习驱动研究中的可复现性,分析障碍与驱动因素。
- 将可复现性分为三个等级:R1(实验可复现性)、R2(数据可复现性)和R3(方法可复现性)。
- 评估检查清单、模型信息表和意识倡议(如 ReproducedPapers.org、可复现性挑战)等工具。
- 评估大型语言模型(LLMs),特别是GPT-4,在高准确率(86.6%)下验证检查清单合规性。
- 分析期刊层面的干预措施,包括代码/数据可用性要求和预注册,以提升研究可信度。
- 研究案例,如ACM TORS期刊,通过发布可复现性论文和预注册支持可复现性。
![Figure 1 : Degrees of reproducibility . Adapted from [ 20 ]](https://ar5iv.labs.arxiv.org/html/2307.10320/assets/degrees.png)
实验结果
研究问题
- RQ1在不同科学领域中,机器学习研究的可复现性当前处于何种水平和程度?
- RQ2哪些主要障碍(如代码和数据不可用或实现敏感性)阻碍了机器学习中的可复现性?
- RQ3检查清单、模型信息表和意识宣传活动等提议的驱动因素在提升机器学习可复现性方面的有效性如何?
- RQ4像GPT-4这样的大型语言模型在多大程度上能够协助验证可复现性检查清单的合规性?
- RQ5期刊和会议如何通过实施政策(如预注册、强制性代码/数据共享)来提升机器学习研究的可复现性?
主要发现
- 机器学习的可复现性危机普遍存在,未公开代码和对训练条件的敏感性是阻碍结果复现的主要障碍。
- R1(实验可复现性)确保在相同条件下输出的一致性,而R3(方法可复现性)则允许在不同实现和数据之间进行泛化。
- 模型信息表通过要求作者记录数据划分和预处理方式,显著降低了数据泄露风险,尤其对非专家机器学习从业者有帮助。
- 如GPT-4等大型语言模型在验证检查清单合规性方面表现出86.6%的准确率,为自动化可复现性验证提供了可扩展的解决方案。
- 意识倡议(如可复现性挑战)以及像 ReproducedPapers.org 这类仓库在促进透明度和教学可复现性方面非常有效。
- ACM TORS等期刊通过支持预注册和发布专门的可复现性研究,正在推进可复现性,从而提升研究可信度。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。