[Paper Review] Machine Learning tools for global PDF fits
This paper presents a machine learning framework within the NNPDF global QCD analysis that uses multi-layer feed-forward neural networks for model-independent parton distribution function (PDF) parametrization, genetic and covariance matrix adaptation algorithms for optimization, and closure testing for systematic validation. The approach achieves robust, statistically sound PDF fits with improved accuracy and uncertainty quantification compared to previous methods.
The use of machine learning algorithms in theoretical and experimental high-energy physics has experienced an impressive progress in recent years, with applications from trigger selection to jet substructure classification and detector simulation among many others. In this contribution, we review the machine learning tools used in the NNPDF family of global QCD analyses. These include multi-layer feed-forward neural networks for the model-independent parametrisation of parton distributions and fragmentation functions, genetic and covariance matrix adaptation algorithms for training and optimisation, and closure testing for the systematic validation of the fitting methodology.
Motivation & Objective
- To develop a model-independent, machine learning-based approach for parametrizing proton parton distribution functions (PDFs) and fragmentation functions (FFs) in global QCD analyses.
- To improve the accuracy and reliability of PDF determinations by employing advanced optimization algorithms such as genetic algorithms and covariance matrix adaptation.
- To systematically validate the fitting methodology using closure testing with pseudo-data generated from known underlying PDF sets.
- To ensure robust uncertainty propagation and statistical consistency in PDF fits through Monte Carlo replica methods and Bayesian reweighting.
- To deliver high-precision PDF sets compatible with the LHAPDF standard for use in LHC phenomenology.
Proposed method
- Neural networks are used as universal interpolants to parametrize PDFs and FFs at a low scale (Q₀ ≈ 1 GeV), enabling model-independent, flexible functional forms.
- Genetic algorithms and covariance matrix adaptation evolution strategies are employed to optimize the neural network parameters by minimizing a χ² figure of merit across experimental data.
- A look-back cross-validation stopping criterion is applied, where training halts at the iteration minimizing the χ² on a validation dataset to prevent overfitting.
- Monte Carlo replicas are used to propagate experimental uncertainties into PDF uncertainties, enabling statistical error estimation and propagation.
- Closure testing is performed by generating pseudo-data from known PDF sets (e.g., MMHT14, CT14), then re-fitting to verify that central values and uncertainties are correctly reproduced.
- Bayesian reweighting is used as a cross-check to confirm that the resulting PDF uncertainties admit a consistent statistical interpretation.
Experimental results
Research questions
- RQ1Can multi-layer feed-forward neural networks effectively serve as model-independent parametrizations for parton distribution functions in global QCD fits?
- RQ2How do genetic and covariance matrix adaptation algorithms compare in optimizing complex, high-dimensional PDF parameter spaces?
- RQ3To what extent does the look-back cross-validation stopping criterion prevent overfitting in neural network-based PDF fits?
- RQ4Can closure testing with pseudo-data demonstrate the statistical robustness and reliability of the NNPDF fitting methodology?
- RQ5Do the PDF uncertainties derived from Monte Carlo replicas and Bayesian reweighting yield consistent and well-calibrated results in closure tests?
Key findings
- The NNPDF3.0 fits using genetic algorithms show improved performance over the NNPDF2.3 fits, particularly in achieving lower χ² values in closure tests at Level 0.
- In Level 2 closure tests, the distribution of single replica fits matches a Gaussian distribution, confirming that central values fluctuate consistently with the quoted PDF uncertainties.
- The look-back cross-validation stopping criterion successfully prevents overfitting by identifying the optimal training iteration where validation χ² reaches a minimum.
- At Level 0 closure tests, the χ² value becomes arbitrarily small because the input PDF set (e.g., MMHT14) yields a perfect fit, confirming the method's ability to recover known solutions.
- Bayesian reweighting successfully reproduces the fit results, providing a cross-validation that the PDF uncertainties have a robust statistical interpretation.
- The NNPDF framework successfully delivers PDF sets compatible with the LHAPDF standard, enabling integration into LHC analysis pipelines.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.