Skip to main content
QUICK REVIEW

[论文解读] ASCR/HEP Exascale Requirements Review Report

Salman Habib, R. Roser|arXiv (Cornell University)|Mar 30, 2016
Advanced Data Storage Technologies参考文献 12被引用 4
一句话总结

本报告概述了至2025年美国高能物理(HEP)实现百亿亿次计算所需的要求,指出高亮度LHC(HL-LHC)、深部地下中微子实验(DUNE)和大型合成巡天望远镜(LSST)推动了数据量和计算需求增长约100倍。报告提出建立HEP-ASCR协作框架,以扩展高性能计算(HPC)、数据管理和网络基础设施,重点聚焦代码重构、边缘服务和混合资源模型,以应对百亿亿次规模的工作负载。

ABSTRACT

This draft report summarizes and details the findings, results, and recommendations derived from the ASCR/HEP Exascale Requirements Review meeting held in June, 2015. The main conclusions are as follows. 1) Larger, more capable computing and data facilities are needed to support HEP science goals in all three frontiers: Energy, Intensity, and Cosmic. The expected scale of the demand at the 2025 timescale is at least two orders of magnitude -- and in some cases greater -- than that available currently. 2) The growth rate of data produced by simulations is overwhelming the current ability, of both facilities and researchers, to store and analyze it. Additional resources and new techniques for data analysis are urgently needed. 3) Data rates and volumes from HEP experimental facilities are also straining the ability to store and analyze large and complex data volumes. Appropriately configured leadership-class facilities can play a transformational role in enabling scientific discovery from these datasets. 4) A close integration of HPC simulation and data analysis will aid greatly in interpreting results from HEP experiments. Such an integration will minimize data movement and facilitate interdependent workflows. 5) Long-range planning between HEP and ASCR will be required to meet HEP's research needs. To best use ASCR HPC resources the experimental HEP program needs a) an established long-term plan for access to ASCR computational and data resources, b) an ability to map workflows onto HPC resources, c) the ability for ASCR facilities to accommodate workflows run by collaborations that can have thousands of individual members, d) to transition codes to the next-generation HPC platforms that will be available at ASCR facilities, e) to build up and train a workforce capable of developing and using simulations and analysis to support HEP scientific research on next-generation systems.

研究动机与目标

  • 应对下一代HEP实验(如HL-LHC、DUNE和LSST)带来的计算与数据管理需求持续增长。
  • 识别支持2025年百亿亿次规模模拟与数据处理所需的HPC、数据和网络基础设施关键需求。
  • 促进HEP与ASCR社区之间的协作,共同开发面向百亿亿次计算准备的基础设施与软件。
  • 通过统一的工作流与监控系统,实现对混合计算资源(包括HPC、云和HTC)的高效利用。
  • 确保HEP软件与数据处理管道针对现代架构(包括加速器支持和I/O感知设计)进行优化。

提出的方法

  • 与领先的HEP与ASCR专家合作,开展系统性需求评审,以评估未来HPC需求。
  • 分析HL-LHC、DUNE和LSST带来的数据增长(最高达当前水平的200倍)与处理需求。
  • 评估当前HEP工作负载:事件模拟(1.5 MB/事件,约50秒/重建)、数据重建(1 GB/文件,约9小时)和分析。
  • 提出对HEP代码库进行重构,以充分利用HPC与加速器架构,包括GPU和多核并行处理。
  • 倡导采用边缘计算、优化数据缓存以及增强网络(如ESnet、Internet2)以减少I/O瓶颈。
  • 建议采用混合资源模型,基于成本、可用性与任务关键性,实现HPC、云和HTC之间的自动作业调度。

实验结果

研究问题

  • RQ1在HL-LHC、DUNE和LSST的背景下,为支持2025年前HEP科学,需要何种规模的计算与数据基础设施?
  • RQ2如何重构HEP工作负载,以高效利用包括加速器和高吞吐系统在内的百亿亿次HPC架构?
  • RQ3混合计算模型(结合HPC、云和HTC)在应对动态HEP工作负载需求方面应扮演何种角色?
  • RQ4如何优化数据I/O与网络性能,以防止模拟与重建工作负载中的性能瓶颈?
  • RQ5为协同设计与部署面向百亿亿次计算的基础设施与数据服务,HEP与ASCR之间需要何种协作框架?

主要发现

  • 仅HL-LHC就将需要接近10倍的处理能力提升与10倍的数据量增长,符合硬件发展趋势,但要求新的基础设施支持。
  • 下一代实验的数据产量可能达到当前水平的200倍,因此需要磁带存储与高容量数据管理系统。
  • 在当前硬件上,1 GB模拟数据文件的事件重建耗时约9小时,凸显I/O与处理瓶颈。
  • HEP中HPC的平均利用率超过95%,表明效率较高,但也凸显了对可扩展、I/O优化系统的需求。
  • 通过自动化的、基于成本的作业调度,将HPC、云与HTC资源集成,对实现弹性与高性能至关重要。
  • 对HEP软件进行重构,以支持并行处理、内存扩展与加速器支持,是实现百亿亿次计算准备的关键。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。