Skip to main content
QUICK REVIEW

[Paper Review] Snowmass 2013 Computing Frontier Storage and Data Management

M. E Butler, R. Mount|arXiv (Cornell University)|Nov 18, 2013
Distributed and Parallel Computing Systems4 references3 citations
TL;DR

This paper proposes a shift toward virtual data and optimized, cost-effective storage hierarchies—leveraging tape, disk, and solid-state storage—to reduce the operational burden of data-intensive HEP experiments. It advocates for a unified, low-cost data management framework inspired by LHC-scale systems but adapted for broader scientific use, with virtual data enabling dynamic, demand-driven data instantiation and long-term preservation.

ABSTRACT

The data storage and data management needs are summarized for the energy frontier, intensity frontier, cosmic frontier, lattice field theory, perturbative QCD and accelerator science. The outlook for data storage technologies and costs is then outlined, followed by a summary of the current state of data, software and physics analysis capability preservation. The HEP outlook is summarized, pointing out where future data volumes may strain against what is technologically and financially feasible. Finally recommendations for areas of particular attention and action are made.

Motivation & Objective

  • Address the escalating cost and complexity of storing and managing petabyte-scale data in high-energy physics (HEP) experiments.
  • Reduce operational expenses of distributed data management systems, currently costing ~$4.5M/year for US-ATLAS alone.
  • Enable long-term data preservation and open access by developing shared, standardized infrastructure across HEP frontiers.
  • Promote interoperability and reuse of HEP data management technologies beyond HEP through collaboration with broader science and industry.
  • Design flexible, adaptable computing models that respond to evolving storage cost dynamics, especially slowing disk capacity growth.

Proposed method

  • Adopt a virtual data model where data products exist only as executable recipes until needed, minimizing persistent storage.
  • Implement dynamic instantiation and replication of data based on anticipated demand and cost trade-offs between storage and re-creation.
  • Leverage existing provenance and workflow systems (e.g., in ATLAS and CMS) to record rigorous data lineage for virtual data support.
  • Integrate xrootd and ROOT-based persistency with scalable, distributed storage systems for efficient data access and transfer.
  • Optimize storage tiering using tape (low cost, high latency), rotating disk (balanced), and solid-state storage (low latency, high cost) based on access patterns.
  • Monitor and adapt to evolving hardware trends—especially slowing disk capacity growth and declining SSD costs—to maintain cost efficiency.

Experimental results

Research questions

  • RQ1How can HEP experiments reduce the operational cost of distributed data management without sacrificing data availability or physics analysis capability?
  • RQ2To what extent can virtual data models—where data is instantiated on-demand—replace persistent storage of derived and simulated data?
  • RQ3What role can shared, cross-disciplinary infrastructure play in enabling long-term data preservation and open science in HEP and beyond?
  • RQ4How can HEP leverage existing software (e.g., ROOT, xrootd) and provenance systems to support scalable, agile data access and virtual data implementation?
  • RQ5What are the implications of slowing disk capacity/cost improvements for future HEP computing and storage architectures?

Key findings

  • The cost of storing and analyzing data is a major fraction of LHC experiment operations, with ATLAS and CMS spending ~$4.5M/year on M&O for distributed data management alone.
  • LHC experiments currently store and treat all persistent data equally, but flagging a large fraction for tape-only storage could significantly reduce costs.
  • Rotating disk storage capacity/cost improvements are expected to slow markedly, potentially disrupting current HEP computing models.
  • Solid-state storage is expected to become relatively cheaper (by ~3x) and more viable for random or sparse access, but requires application-aware caching to be effective.
  • The virtual data concept—where data exists only as recipes until needed—could enable dynamic adaptation to changing storage and CPU cost ratios.
  • Provenance recording and existing workflow systems in LHC experiments already contain much of the infrastructure needed to support virtual data and long-term data preservation.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.