Skip to main content
QUICK REVIEW

[Paper Review] A compression algorithm for the combination of PDF sets

Stefano Carrazza, José I. Latorre|arXiv (Cornell University)|Apr 24, 2015
Particle physics theoretical and experimental studies58 references4 citations
TL;DR

This paper proposes a compression algorithm to combine multiple global PDF sets (NNPDF3.0, CT14, MMHT14) into a single, compact Monte Carlo representation—CMC-PDF—using a statistical framework based on Monte Carlo replicas and a genetic algorithm optimizer. The method reduces 1000 replicas to ~100 while preserving accuracy in LHC cross-sections and parton luminosities, enabling efficient uncertainty propagation in LHC phenomenology.

ABSTRACT

The current PDF4LHC recommendation to estimate uncertainties due to parton distribution functions (PDFs) in theoretical predictions for LHC processes involves the combination of separate predictions computed using PDF sets from different groups, each of which comprises a relatively large number of either Hessian eigenvectors or Monte Carlo (MC) replicas. While many fixed-order and parton shower programs allow the evaluation of PDF uncertainties for a single PDF set at no additional CPU cost, this feature is not universal, and moreover the a posteriori combination of the predictions using at least three different PDF sets is still required. In this work, we present a strategy for the statistical combination of individual PDF sets, based on the MC representation of Hessian sets, followed by a compression algorithm for the reduction of the number of MC replicas. We illustrate our strategy with the combination and compression of the recent NNPDF3.0, CT14 and MMHT14 NNLO PDF sets. The resulting Compressed Monte Carlo PDF (CMC-PDF) sets are validated at the level of parton luminosities and LHC inclusive cross-sections and differential distributions. We determine that around 100 replicas provide an adequate representation of the probability distribution for the original combined PDF set, suitable for general applications to LHC phenomenology.

Motivation & Objective

  • To address the challenge of combining multiple PDF sets from different groups for robust PDF uncertainty estimation in LHC physics.
  • To develop a statistically sound, user-friendly method for combining Hessian and Monte Carlo PDF sets into a single, compressed Monte Carlo representation.
  • To reduce the computational cost of PDF uncertainty propagation while preserving accuracy in LHC cross-sections and differential distributions.
  • To provide a practical, accessible tool for the LHC community to use combined PDF uncertainties without relying on envelope methods.

Proposed method

  • Transform Hessian PDF sets (NNPDF3.0, CT14, MMHT14) into Monte Carlo replicas using the Watt-Thorne method.
  • Combine replicas from each PDF set into a single, unified Monte Carlo PDF set with equal weighting per group.
  • Apply a genetic algorithm-based compression to reduce the number of replicas from 1000 to ~100 while minimizing statistical error.
  • Define an error function based on Kolmogorov-Smirnov distances and moment matching to preserve PDF correlations and distributions.
  • Use LHAPDF6 and ROOT for I/O and validation, with a custom genetic algorithm to optimize replica selection.
  • Validate the compressed sets using parton luminosities and LHC cross-sections across multiple benchmark processes.

Experimental results

Research questions

  • RQ1Can a statistically robust, unified Monte Carlo PDF set be constructed from multiple independent PDF sets (NNPDF3.0, CT14, MMHT14) with consistent uncertainty treatment?
  • RQ2What is the minimum number of replicas needed to accurately represent the combined PDF uncertainty for LHC phenomenology?
  • RQ3How well does the compressed Monte Carlo PDF set reproduce the parton luminosities and cross-sections of the original combined set?
  • RQ4Can the compression algorithm preserve higher-order moments and correlations between PDFs across different $x$ and $Q^2$ regions?

Key findings

  • The CMC-PDF sets, compressed to 100 replicas, accurately reproduce the parton luminosities and inclusive cross-sections of the original combined PDF set.
  • The compression algorithm maintains agreement with the original combined set at the sub-percent level for key LHC processes, including $W$, $Z$, and Higgs production.
  • The Kolmogorov-Smirnov distance and moment-matching criteria ensure that the compressed replicas preserve the statistical properties of the full set.
  • The method reduces the number of replicas from 1000 to 100 with minimal loss of accuracy, making it computationally efficient for widespread use.
  • The CMC-PDF sets are validated across multiple benchmark processes, showing consistency with both individual PDF sets and the PDF4LHC envelope method.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.