[论文解读] Standing Together for Reproducibility in Large-Scale Computing: Report on reproducibility@XSEDE
本文报告了reproducibility@XSEDE研讨会的成果,倡导通过机构和个体责任来确保大规模计算研究的可重现性。论文提出了可操作的策略——如改进文档记录、系统级 provenance 追踪、协作最佳实践以及网关开发——并强调设立年度奖项以表彰可重现研究卓越成果的潜在可能性。
This is the final report on reproducibility@xsede, a one-day workshop held in conjunction with XSEDE14, the annual conference of the Extreme Science and Engineering Discovery Environment (XSEDE). The workshop's discussion-oriented agenda focused on reproducibility in large-scale computational research. Two important themes capture the spirit of the workshop submissions and discussions: (1) organizational stakeholders, especially supercomputer centers, are in a unique position to promote, enable, and support reproducible research; and (2) individual researchers should conduct each experiment as though someone will replicate that experiment. Participants documented numerous issues, questions, technologies, practices, and potentially promising initiatives emerging from the discussion, but also highlighted four areas of particular interest to XSEDE: (1) documentation and training that promotes reproducible research; (2) system-level tools that provide build- and run-time information at the level of the individual job; (3) the need to model best practices in research collaborations involving XSEDE staff; and (4) continued work on gateways and related technologies. In addition, an intriguing question emerged from the day's interactions: would there be value in establishing an annual award for excellence in reproducible research?
研究动机与目标
- 应对大规模计算研究中可重现性挑战日益加剧的问题,因为研究结果往往难以验证或复现。
- 明确超级计算中心(如XSEDE)在推动和促进可重现研究实践方面的独特作用。
- 鼓励个体研究人员在设计实验时保持完全透明,如同预期会被他人复现一般。
- 突出将可重现性融入研究生命周期所需的关键技术与文化转变。
- 探讨机构认可机制的可行性,例如设立年度奖项以表彰可重现研究的卓越成就。
提出的方法
- 在XSEDE14会议期间举办了一天的讨论式研讨会,收集来自多样化研究人员和利益相关者的见解。
- 收集了关于大规模计算环境中可重现性相关挑战、工具和最佳实践的反馈。
- 提出了可在单个作业级别捕获构建和运行时 provenance 的系统级工具,以提升透明度。
- 强调了文档记录和培训项目在将可重现性嵌入研究工作流程中的重要性。
- 倡导在XSEDE工作人员与外部研究人员的合作中,以最佳实践为模型开展研究协作。
- 探讨了科学网关的开发与增强,作为支持可重现研究实践的平台。
实验结果
研究问题
- RQ1超级计算中心(如XSEDE)如何能积极促进并支持大规模计算项目中的可重现研究?
- RQ2为确保计算实验能被他人有意义地复现,需要哪些技术和程序上的改进?
- RQ3标准化文档记录和培训在多大程度上可以降低高性能计算中可重现性的障碍?
- RQ4系统级 provenance 追踪工具在提升计算工作流的透明度和可审计性方面发挥什么作用?
- RQ5设立年度可重现研究卓越奖项是否有助于在科学界制度化并激励最佳实践?
主要发现
- 组织利益相关者,尤其是超级计算中心,在通过基础设施、政策和支持推动可重现性方面具有独特优势。
- 研究人员应采取透明的思维模式,将复现作为核心原则来设计实验。
- 能够捕获作业级别构建和运行时 provenance 的系统级工具,对于追踪计算工作流并确保可重现性至关重要。
- 改进的文档记录和培训是广泛采用可重现研究实践的关键推动因素。
- 在XSEDE支持的合作中树立最佳实践范例,可为其他研究机构提供模板。
- 提出设立年度可重现研究卓越奖项,作为认可和激励高质量、透明研究的潜在机制。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。