Skip to main content
QUICK REVIEW

[Paper Review] Report of the Snowmass 2013 Computing Frontier Working Group on Distributed Computing and Facility Infrastructures

K. Bloom, R. Gerber|arXiv (Cornell University)|Nov 9, 2013
Distributed and Parallel Computing Systems1 references3 citations
TL;DR

This report evaluates the state and future of distributed computing and facility infrastructures for high-energy physics (HEP), emphasizing the critical role of the Worldwide LHC Computing Grid (WLCG) in supporting LHC experiments through a tiered, globally distributed HTC (high-throughput computing) model. It advocates for continued investment in WLCG, integration of HPC and HTC resources, and adaptation of software to multi-core and GPU architectures to meet growing data and simulation demands from the Energy, Intensity, and Cosmic Frontiers.

ABSTRACT

This is the report of the Snowmass 2013 Computing Frontier Working Group on Distributed Computing and Facility Infrastructures.

Motivation & Objective

  • To assess the current and projected computing infrastructure needs of the U.S. high-energy physics (HEP) community through 2017.
  • To evaluate the role of the Worldwide LHC Computing Grid (WLCG) as the primary distributed computing infrastructure for LHC experiments.
  • To examine the integration of HPC and HTC paradigms in national computing centers and their relevance to future HEP research.
  • To identify challenges in scaling computing resources for increased LHC luminosity, event complexity, and data rates.
  • To recommend strategies for software evolution, resource diversification, and sustained funding to maintain HEP’s computational readiness.

Proposed method

  • Analyzes existing HEP computing workloads, distinguishing between HTC (embarrassingly parallel) and HPC (tightly coupled, high-performance) computing models.
  • Evaluates the WLCG’s tiered architecture—Tier-0 at CERN, Tier-1 sites for data reprocessing and archiving, and Tier-2 sites for physics analysis and simulation.
  • Reviews the role of national computing centers (e.g., NERSC) in supporting both HTC and HPC workloads, including software and middleware development.
  • Assesses the impact of increasing LHC data rates (from 300 Hz to 1 kHz) and pileup (20 to 25 interactions per event) on infrastructure scalability.
  • Proposes evolutionary upgrades to WLCG and integration of opportunistic resources, including commercial clouds and national HPC centers.
  • Recommends software modernization for multi-core and GPU-based parallelism to leverage emerging hardware capabilities.

Experimental results

Research questions

  • RQ1How can the WLCG infrastructure evolve to meet the anticipated increase in LHC data rates and event complexity through 2017 and beyond?
  • RQ2To what extent can national HPC centers support HEP workloads traditionally handled by HTC systems, and what changes are needed to enable this?
  • RQ3What are the key technical and operational challenges in scaling distributed computing infrastructures for future HEP experiments across the Energy, Intensity, and Cosmic Frontiers?
  • RQ4How can HEP software be adapted to efficiently utilize massively multi-core and GPU-accelerated architectures?
  • RQ5What role should coordinated resource sharing and training play in enabling Intensity Frontier experiments to access available computing facilities?

Key findings

  • The WLCG, with 2 MHS06 of CPU and 190 PB of disk, successfully supported 300,000 cores continuously in 2012, delivering 2.6 billion CPU hours for LHC data processing and analysis.
  • The WLCG’s tiered, globally distributed HTC model remains well-suited for LHC’s embarrassingly parallel workloads and is expected to remain the primary computational resource for LHC experiments.
  • HEP’s demand for HPC resources is projected to exceed supply by a factor of four by 2017, highlighting a critical shortage in national HPC centers.
  • National centers such as NERSC are increasingly capable of supporting both HTC and HPC workloads, but HEP experiments must proactively engage with these centers to access resources.
  • Software must evolve to exploit fine-grained (e.g., GPU) and coarse-grained (e.g., MPI) parallelism to fully utilize emerging multi-core and accelerated architectures.
  • Opportunistic use of commercial clouds, university clusters, and national HPC centers is essential to diversify computing architectures and ensure long-term sustainability amid rising data demands.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.