[Paper Review] Converging High-Throughput and High-Performance Computing: A Case Study.
This paper presents a case study of ATLAS integrating the Titan leadership-class supercomputer with its existing high-throughput computing infrastructure, achieving sustained production workloads of ~51 million core-hours annually. It evaluates design and operational strategies for scalable supercomputer use, introduces a next-generation PanDA executor for advanced workloads, and establishes a framework for integrating experimental systems with production supercomputing platforms.
The computing systems used by LHC experiments has historically consisted of the federation of hundreds to thousands of distributed resources, ranging from small to mid-size resource. In spite of the impressive scale of the existing distributed computing solutions, the federation of small to mid-size resources will be insufficient to meet projected future demands. This paper is a case study of how the ATLAS experiment has embraced Titan -- a DOE leadership facility in conjunction with traditional distributed high-throughput computing to reach sustained production scales of approximately 51M core-hours a years. The three main contributions of this paper are: (i) a critical evaluation of design and operational considerations to support the sustained, scalable and production usage of Titan; (ii) a preliminary characterization of a next generation executor for PanDA to support new workloads and advanced execution modes; and (iii) early lessons for how current and future experimental and observational systems can be integrated with production supercomputers and other platforms in a general and extensible manner.
Motivation & Objective
- To evaluate the design and operational challenges of using leadership-class supercomputers like Titan for sustained, large-scale production workloads in high-energy physics.
- To develop and characterize a next-generation executor for the PanDA workload management system to support emerging workloads and advanced execution modes.
- To establish a general, extensible framework for integrating experimental and observational systems with production supercomputing platforms.
- To assess the feasibility and performance of combining high-throughput and high-performance computing in a production environment.
Proposed method
- Integration of the ATLAS workload management system (PanDA) with the Titan leadership-class supercomputer at ORNL.
- Adaptation of PanDA's job execution and scheduling components to support Titan's architecture and job submission workflows.
- Design and evaluation of a new executor module in PanDA to handle diverse workloads, including those requiring low-latency or high-concurrency execution.
- Performance characterization of the hybrid HPC-HTC environment using real production workloads from the ATLAS experiment.
- Operational analysis of job submission, monitoring, and failure recovery patterns on Titan in a sustained production setting.
- Development of an extensible software stack enabling interoperability between high-throughput and high-performance computing platforms.
Experimental results
Research questions
- RQ1How can leadership-class supercomputers like Titan be effectively and sustainably integrated into existing high-throughput computing infrastructures for large-scale scientific workloads?
- RQ2What architectural and operational changes are required to enable production-level usage of supercomputers within a high-throughput computing framework like PanDA?
- RQ3How can a next-generation PanDA executor be designed to support diverse and advanced execution modes required by modern scientific workloads?
- RQ4What are the key challenges and performance characteristics when running sustained, large-scale workloads on a leadership-class supercomputer in a production HEP environment?
- RQ5What general patterns and abstractions can be derived to enable seamless integration of experimental systems with heterogeneous HPC and HTC platforms?
Key findings
- The integration of Titan with ATLAS's high-throughput computing infrastructure enabled sustained production workloads of approximately 51 million core-hours per year.
- The design and operational considerations for using Titan in production were successfully evaluated, demonstrating scalability and reliability over extended periods.
- A next-generation executor for PanDA was developed and characterized, enabling support for new workload patterns and execution modes.
- The hybrid HPC-HTC environment achieved high utilization and effective job management across diverse computing resources.
- The study established a general and extensible framework for integrating experimental systems with production supercomputing platforms, applicable to future HEP and observational science workloads.
- Early operational lessons were identified, including the importance of automated monitoring, fault tolerance, and dynamic resource allocation in large-scale HPC-HTC workflows.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.