Skip to main content
QUICK REVIEW

[Paper Review] ATLAS Data Challenge 1

G. Poulard|ArXiv.org|Jun 12, 2003
Distributed and Parallel Computing Systems3 citations
TL;DR

This paper presents ATLAS Data Challenge 1 (DC1), a large-scale distributed computing exercise to validate the ATLAS experiment's computing model, software suite, and data model ahead of the Large Hadron Collider's operation. It describes the successful production of over 10 million physics events and 30 million single-particle events across 39 institutes in 18 countries using grid middleware, achieving 71,000 CPU-days and generating 30 terabytes of data in 35,000 partitions, demonstrating the feasibility of global, collaborative Monte Carlo simulation under a distributed computing framework.

ABSTRACT

In 2002 the ATLAS experiment started a series of Data Challenges (DC) of which the goals are the validation of the Computing Model, of the complete software suite, of the data model, and to ensure the correctness of the technical choices to be made. A major feature of the first Data Challenge (DC1) was the preparation and the deployment of the software required for the production of large event samples for the High Level Trigger (HLT) and physics communities, and the production of those samples as a world-wide distributed activity. The first phase of DC1 was run during summer 2002, and involved 39 institutes in 18 countries. More than 10 million physics events and 30 million single particle events were fully simulated. Over a period of about 40 calendar days 71000 CPU-days were used producing 30 Tbytes of data in about 35000 partitions. In the second phase the next processing step was performed with the participation of 56 institutes in 21 countries (~ 4000 processors used in parallel). The basic elements of the ATLAS Monte Carlo production system are described. We also present how the software suite was validated and the participating sites were certified. These productions were already partly performed by using different flavours of Grid middleware at ~ 20 sites.

Motivation & Objective

  • To validate the ATLAS computing model, software suite, and data model ahead of the Large Hadron Collider's operation.
  • To test the scalability and reliability of a worldwide distributed computing infrastructure for high-energy physics.
  • To certify participating sites and ensure interoperability across diverse computing environments using grid middleware.
  • To produce and process large-scale Monte Carlo event samples for the High Level Trigger and physics communities.
  • To establish a foundation for future data challenges and production workflows in the ATLAS experiment.

Proposed method

  • The first phase of DC1 involved 39 institutes in 18 countries deploying and executing the ATLAS Monte Carlo production software suite on distributed computing resources.
  • The production used various flavors of grid middleware at approximately 20 sites to enable interoperability and workload distribution.
  • A total of 71,000 CPU-days were consumed over 40 calendar days to simulate 10 million physics events and 30 million single-particle events.
  • The data were partitioned into approximately 35,000 files to support scalable processing and storage in a distributed environment.
  • The second phase involved 56 institutes in 21 countries, using around 4,000 processors in parallel for subsequent processing steps.
  • Site certification was performed based on successful execution of production tasks and compliance with ATLAS software and data standards.

Experimental results

Research questions

  • RQ1Can a globally distributed computing infrastructure successfully produce and manage large-scale Monte Carlo event samples for the ATLAS experiment?
  • RQ2How effective is the ATLAS software suite in a heterogeneous, multi-institutional computing environment?
  • RQ3To what extent can different grid middleware implementations interoperate in a high-throughput physics data production workflow?
  • RQ4What are the performance and scalability characteristics of the ATLAS Monte Carlo production system under real-world distributed conditions?
  • RQ5How can site certification and software validation be systematically achieved across a large number of international institutions?

Key findings

  • The ATLAS Data Challenge 1 successfully produced over 10 million physics events and 30 million single-particle events using a distributed computing model.
  • A total of 71,000 CPU-days were consumed over 40 calendar days, resulting in 30 terabytes of data stored across approximately 35,000 partitions.
  • The production was executed across 39 institutes in 18 countries, demonstrating the feasibility of international collaboration in high-energy physics computing.
  • The second phase involved 56 institutes in 21 countries, with around 4,000 processors used in parallel, confirming scalability of the workflow.
  • The software suite was successfully validated, and participating sites were certified based on consistent and correct execution of production tasks.
  • The use of multiple grid middleware flavors at ~20 sites confirmed interoperability and robustness of the distributed computing framework.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.