Skip to main content
QUICK REVIEW

[论文解读] Converging High-Throughput and High-Performance Computing: A Case Study.

Alessio Angius, D. Oleynik|arXiv (Cornell University)|Apr 4, 2017
Distributed and Parallel Computing Systems被引用 3
一句话总结

本文展示了ATLAS将Titan领导级超算系统与现有高吞吐量计算基础设施集成的案例研究,实现了每年约5100万核时的持续生产工作负载。研究评估了可扩展超算使用的架构与运行策略,提出了下一代PanDA执行器以支持高级工作负载,并建立了一个将实验系统与生产级超算平台集成的框架。

ABSTRACT

The computing systems used by LHC experiments has historically consisted of the federation of hundreds to thousands of distributed resources, ranging from small to mid-size resource. In spite of the impressive scale of the existing distributed computing solutions, the federation of small to mid-size resources will be insufficient to meet projected future demands. This paper is a case study of how the ATLAS experiment has embraced Titan -- a DOE leadership facility in conjunction with traditional distributed high-throughput computing to reach sustained production scales of approximately 51M core-hours a years. The three main contributions of this paper are: (i) a critical evaluation of design and operational considerations to support the sustained, scalable and production usage of Titan; (ii) a preliminary characterization of a next generation executor for PanDA to support new workloads and advanced execution modes; and (iii) early lessons for how current and future experimental and observational systems can be integrated with production supercomputers and other platforms in a general and extensible manner.

研究动机与目标

  • 评估在高能物理领域中,使用Titan等领导级超算系统实现持续、大规模生产工作负载的架构与运行挑战。
  • 开发并表征下一代PanDA工作负载管理系统执行器,以支持新兴工作负载和高级执行模式。
  • 建立一个通用且可扩展的框架,用于将实验与观测系统与生产级超算平台集成。
  • 评估在生产环境中结合高吞吐量计算与高性能计算的可行性与性能。

提出的方法

  • 将ATLAS工作负载管理系统(PanDA)与ORNL的Titan领导级超算系统集成。
  • 适配PanDA的作业执行与调度组件,以支持Titan的架构和作业提交工作流。
  • 设计并评估PanDA中新的执行器模块,以处理多样化的作业负载,包括需要低延迟或高并发执行的作业。
  • 使用来自ATLAS实验的真实生产工作负载,对混合HPC-HTC环境的性能进行表征。
  • 在持续生产环境中,对Titan上的作业提交、监控与故障恢复模式进行运行分析。
  • 开发一个可扩展的软件堆栈,以实现高吞吐量计算与高性能计算平台之间的互操作性。

实验结果

研究问题

  • RQ1如何有效且可持续地将Titan等领导级超算系统集成到现有高吞吐量计算基础设施中,以支持大规模科学工作负载?
  • RQ2在高吞吐量计算框架(如PanDA)中实现超算系统生产级使用,需要哪些架构与运行变更?
  • RQ3如何设计下一代PanDA执行器,以支持现代科学工作负载所需的多样化与高级执行模式?
  • RQ4在生产HEP环境中,在领导级超算系统上持续运行大规模工作负载时,面临的主要挑战与性能特征是什么?
  • RQ5可以提炼出哪些通用模式与抽象机制,以实现实验系统与异构HPC与HTC平台之间的无缝集成?

主要发现

  • ATLAS将Titan与高吞吐量计算基础设施集成,成功实现了每年约5100万核时的持续生产工作负载。
  • 对Titan在生产环境中使用的架构与运行考量进行了成功评估,证明了其在长时间运行中的可扩展性与可靠性。
  • 开发并表征了下一代PanDA执行器,支持新型工作负载模式与执行模式。
  • 混合HPC-HTC环境实现了高资源利用率,并在多样化计算资源上实现了有效的作业管理。
  • 本研究建立了一个通用且可扩展的框架,用于将实验系统与生产级超算平台集成,适用于未来的HEP与观测科学工作负载。
  • 早期运行经验表明,自动化监控、容错能力与动态资源分配在大规模HPC-HTC工作流中至关重要。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。