[Paper Review] Improving Reproducibility in Machine Learning Research (A Report from the NeurIPS 2019 Reproducibility Program)
This paper documents NeurIPS 2019's reproducibility program, detailing code submission policy, reproducibility challenge, and ML reproducibility checklist, and reports community uptake and preliminary effects on review quality.
One of the challenges in machine learning research is to ensure that presented and published results are sound and reliable. Reproducibility, that is obtaining similar results as presented in a paper or talk, using the same code and data (when available), is a necessary step to verify the reliability of research findings. Reproducibility is also an important step to promote open and accessible research, thereby allowing the scientific community to quickly integrate new findings and convert ideas to practice. Reproducibility also promotes the use of robust experimental workflows, which potentially reduce unintentional errors. In 2019, the Neural Information Processing Systems (NeurIPS) conference, the premier international conference for research in machine learning, introduced a reproducibility program, designed to improve the standards across the community for how we conduct, communicate, and evaluate machine learning research. The program contained three components: a code submission policy, a community-wide reproducibility challenge, and the inclusion of the Machine Learning Reproducibility checklist as part of the paper submission process. In this paper, we describe each of these components, how it was deployed, as well as what we were able to learn from this initiative.
Motivation & Objective
- Promote transparency by encouraging sharing of code, data, and artifacts alongside ML papers.
- Assess the impact of reproducibility practices on paper quality and reviewer experience.
- Explore community engagement with reproducibility challenges and checklists.
- Provide guidelines to inform broader adoption of reproducibility practices across ML venues.
Proposed method
- Describe three components of NeurIPS 2019 reproducibility program: code submission policy, reproducibility challenge, and ML reproducibility checklist.
- Implement checklist at initial submission and camera-ready phases to analyze answer changes.
- Use OpenReview and public reproducibility reports to foster transparency and replication.
- Analyze reviewer engagement with code and checklist responses and associated paper outcomes.
- Compare code availability and acceptance rates across conferences to contextualize policy effects.
Experimental results
Research questions
- RQ1What are the effects of a code submission policy on reviewer behavior and paper acceptance?
- RQ2Does participation in a reproducibility challenge increase reproduction efforts and transparency?
- RQ3How useful is the ML reproducibility checklist for authors and reviewers, and does it correlate with paper quality?
- RQ4What are broader implications for adopting reproducibility practices in ML venues?
Key findings
- Code submission participation rose to about 75% by camera-ready, with reviewers frequently consulting code when available.
- Reviewers who consulted or had access to code tended to assign higher scores to papers (statistical association observed).
- The reproducibility challenge saw growing participation and reports, with 173 papers claimed for reproduction across 73 institutions in NeurIPS 2019.
- Checklist answers showed that around one-third of reviewers found it useful, and usefulness correlated with higher paper scores and reviewer confidence.
- Overall conference submissions increased (~40%), suggesting no drop in interest due to reproducibility initiatives.
- A significant portion of authors provided code at submission or camera-ready stages, indicating growing openness to artifacts.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.