[论文解读] Roadmap on Electronic Structure Codes in the Exascale Era
本路线图概述了14种主要电子结构代码在面向百亿亿次计算时的当前状态与未来发展方向。它识别了在大规模并行计算和硬件加速器方面扩展所面临的挑战,并详细说明了为实现材料科学、化学和器件物理领域更大规模、更高精度模拟而需优先发展的方向。
Electronic structure calculations have been instrumental in providing many important insights into a range of physical and chemical properties of various molecular and solid-state systems. Their importance to various fields, including materials science, chemical sciences, computational chemistry and device physics, is underscored by the large fraction of available public supercomputing resources devoted to these calculations. As we enter the exascale era, exciting new opportunities to increase simulation numbers, sizes, and accuracies present themselves. In order to realize these promises, the community of electronic structure software developers will however first have to tackle a number of challenges pertaining to the efficient use of new architectures that will rely heavily on massive parallelism and hardware accelerators. This roadmap provides a broad overview of the state-of-the-art in electronic structure calculations and of the various new directions being pursued by the community. It covers 14 electronic structure codes, presenting their current status, their development priorities over the next five years, and their plans towards tackling the challenges and leveraging the opportunities presented by the advent of exascale computing.
研究动机与目标
- 应对材料科学、化学和器件物理领域电子结构计算日益增长的计算需求。
- 识别百亿亿次系统带来的架构与算法挑战,特别是大规模并行计算和异构加速器。
- 指导电子结构软件社区调整开发路线图,以在下一代超级计算机上实现最高效率和可扩展性。
- 促进14个主要代码之间的协调合作,确保互操作性、性能可移植性以及在百亿亿次时代的长期可持续性。
提出的方法
- 调查了14种领先的电子结构代码,以评估其当前能力与未来开发轨迹。
- 评估每种代码在百亿亿次架构上实现高并行效率的策略,包括基于任务和数据驱动的并行计算。
- 分析利用GPU和专用AI加速器等硬件加速器的迁移路径。
- 探索软件栈创新,包括性能可移植性框架(如Kokkos、RAJA)以及用于负载均衡的运行时系统。
- 回顾算法进展,如线性标度方法、混合基组方法以及自适应精度,以降低计算成本。
- 强调模块化、可扩展的软件设计对于支持不断演进的硬件和科学需求的必要性。
实验结果
研究问题
- RQ1如何高效地将电子结构代码扩展至具备大规模并行计算和异构加速器的百亿亿次架构?
- RQ2在多样化的百亿亿次系统上实现性能可移植性的关键软件与算法挑战是什么?
- RQ3在面向百亿亿次时代的背景下,领先电子结构代码中正在浮现的共同开发优先事项与架构策略有哪些?
- RQ4社区如何确保电子结构软件在硬件快速演进背景下的长期可持续性与互操作性?
- RQ5新型编程模型与运行时系统在实现高效且可移植的百亿亿次模拟中发挥什么作用?
主要发现
- 所调查的14种代码正在积极推进向百亿亿次系统的迁移,重点集中在GPU加速和基于任务的并行计算。
- Kokkos和RAJA等性能可移植性框架被广泛采用,以应对硬件异构性并提升可维护性。
- 许多代码正在投资于算法创新,如线性标度方法和混合基组方法,以降低大规模计算的计算成本。
- 一种对模块化、可扩展软件设计的共同重视正在浮现,以支持长期可持续性,并与新兴硬件实现集成。
- 该路线图指出了在软件开发方面开展全社区协调的迫切需求,以避免碎片化并确保百亿亿次资源的高效利用。
- 在支持更大规模系统和更高精度模拟方面已取得显著进展,预计在百亿亿次平台上系统规模和模拟吞吐量将大幅提升。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。