Skip to main content
QUICK REVIEW

[论文解读] ATLAS Data Challenge 1

G. Poulard|ArXiv.org|Jun 12, 2003
Distributed and Parallel Computing Systems被引用 3
一句话总结

本文介绍了ATLAS数据挑战1(DC1),这是一次大规模的分布式计算实验,旨在验证大型强子对撞机(LHC)运行前ATLAS实验的计算模型、软件套件和数据模型。该研究描述了利用网格中间件,在18个国家的39个机构中成功生成超过1000万次物理事件和3000万次单粒子事件,累计消耗71,000个CPU日,生成30TB数据并划分为35,000个数据分区,证明了在全球分布式计算框架下开展大规模协作蒙特卡洛模拟的可行性。

ABSTRACT

In 2002 the ATLAS experiment started a series of Data Challenges (DC) of which the goals are the validation of the Computing Model, of the complete software suite, of the data model, and to ensure the correctness of the technical choices to be made. A major feature of the first Data Challenge (DC1) was the preparation and the deployment of the software required for the production of large event samples for the High Level Trigger (HLT) and physics communities, and the production of those samples as a world-wide distributed activity. The first phase of DC1 was run during summer 2002, and involved 39 institutes in 18 countries. More than 10 million physics events and 30 million single particle events were fully simulated. Over a period of about 40 calendar days 71000 CPU-days were used producing 30 Tbytes of data in about 35000 partitions. In the second phase the next processing step was performed with the participation of 56 institutes in 21 countries (~ 4000 processors used in parallel). The basic elements of the ATLAS Monte Carlo production system are described. We also present how the software suite was validated and the participating sites were certified. These productions were already partly performed by using different flavours of Grid middleware at ~ 20 sites.

研究动机与目标

  • 在大型强子对撞机运行前,验证ATLAS计算模型、软件套件和数据模型。
  • 测试面向高能物理的全球分布式计算基础设施的可扩展性和可靠性。
  • 基于网格中间件,对参与机构进行认证,确保在多样化计算环境下的互操作性。
  • 为高级触发器和物理研究社区生成并处理大规模蒙特卡洛事件样本。
  • 为ATLAS实验未来的数据挑战和生产工作流建立基础。

提出的方法

  • DC1的第一阶段涉及18个国家的39个机构,在分布式计算资源上部署并执行ATLAS蒙特卡洛生产软件套件。
  • 生产过程中在约20个站点使用了多种网格中间件版本,以实现互操作性和工作负载分发。
  • 在40个日历日内共消耗71,000个CPU日,用于模拟1000万次物理事件和3000万次单粒子事件。
  • 数据被划分为约35,000个文件,以支持在分布式环境中的可扩展处理与存储。
  • 第二阶段涉及21个国家的56个机构,使用约4,000个处理器并行执行后续处理步骤。
  • 基于生产任务的成功执行以及对ATLAS软件和数据标准的符合性,完成了站点认证。

实验结果

研究问题

  • RQ1全球分布的计算基础设施能否成功生成并管理ATLAS实验的大规模蒙特卡洛事件样本?
  • RQ2ATLAS软件套件在异构、多机构计算环境中表现如何?
  • RQ3不同网格中间件实现方式在高吞吐量物理数据生产工作流中可互操作的程度如何?
  • RQ4ATLAS蒙特卡洛生产系统在真实世界分布式条件下的性能与可扩展性特征是什么?
  • RQ5如何在大量国际机构中系统性地实现站点认证与软件验证?

主要发现

  • ATLAS数据挑战1成功利用分布式计算模型生成了超过1000万次物理事件和3000万次单粒子事件。
  • 在40个日历日内共消耗71,000个CPU日,生成30TB数据并存储于约35,000个数据分区中。
  • 生产过程在18个国家的39个机构中完成,证明了高能物理计算中国际合作的可行性。
  • 第二阶段涉及21个国家的56个机构,约4,000个处理器并行运行,证实了工作流的可扩展性。
  • 软件套件成功通过验证,参与机构基于生产任务的一致且正确执行获得认证。
  • 在约20个站点使用多种网格中间件版本,证实了分布式计算框架的互操作性与稳健性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。