[Paper Review] Benchmarking Machine Learning Techniques with Di-Higgs Production at the LHC
This paper benchmarks multiple machine learning techniques—ranging from traditional boosted decision trees to deep neural networks and autoencoders—for identifying di-Higgs boson production (hh → bb̄bb̄) at the HL-LHC. Using simulated 14 TeV proton-proton collisions with 3000 fb⁻¹ luminosity, it finds that convolutional neural networks achieve the highest significance of 2.85 ± 0.02, significantly outperforming classical cuts and unsupervised autoencoders, which yield only 0.81 ± 0.01 due to poor signal-background separation in latent space.
Many domains of high energy physics analysis are starting to explore machine learning techniques. Powerful methods can be used to identify and measure rare processes from previously insurmountable backgrounds. One of the most profound Standard Model signatures still to be discovered at the LHC is the pair production of Higgs bosons through the Higgs self-coupling. The small cross section of this process makes detection very difficult even for the decay channel with the largest branching fraction ($hh ightarrow b\bar{b}b\bar{b}$). This paper benchmarks a variety of approaches (boosted decision trees, various neural network architectures, semi-supervised algorithms) against one another to catalog a few of the various techniques available to high energy physicists as the era of the HL-LHC approaches.
Motivation & Objective
- To evaluate and compare diverse machine learning techniques for identifying rare di-Higgs production in the dominant bb̄bb̄ decay channel at the HL-LHC.
- To assess the performance of supervised and unsupervised ML methods under idealized zero-pileup conditions, simulating the HL-LHC environment.
- To identify which architectures are most effective at separating the rare di-Higgs signal from dominant QCD multijet backgrounds.
- To provide a reference for future high-energy physics analyses by cataloging the sensitivity gains of modern ML techniques over traditional kinematic cuts.
Proposed method
- Simulated events for di-Higgs (hh → bb̄bb̄) and QCD multijet backgrounds were generated using MadGraph v2.7.0, showered with Pythia v8.2.44, and reconstructed with Delphes v3.0 for the Phase-II CMS detector.
- A variety of ML models were trained: boosted decision trees, feedforward and convolutional neural networks, random forests, k-means clustering, particle flow networks, and autoencoders with ReLU and sigmoid activations and L2 regularization.
- The autoencoder was trained in an unsupervised manner on QCD events to learn a compressed latent representation, with signal events expected to have higher reconstruction loss.
- Significance was computed as σ = N_signal / √N_background, with yields normalized to 3000 fb⁻¹ of integrated luminosity and statistical uncertainties propagated.
- All models were evaluated under zero-pileup conditions, with results reported for signal and background event counts and significance after selection cuts.
- The study used a simplified event generation with no pileup and minimal hadronic activity cuts to isolate the impact of ML architecture on signal discrimination.
Experimental results
Research questions
- RQ1Which machine learning architecture achieves the highest significance in identifying di-Higgs production (hh → bb̄bb̄) amidst dominant QCD multijet backgrounds under zero-pileup conditions?
- RQ2How do unsupervised methods like autoencoders compare to supervised models in detecting rare di-Higgs events when no labels are used during training?
- RQ3To what extent do deep learning architectures such as convolutional neural networks and particle flow networks outperform traditional kinematic cuts and tree-based models?
- RQ4How do model-specific features and latent space representations influence the separation of signal and background events in high-dimensional phase space?
- RQ5How robust are different ML techniques to reconstruction challenges expected in the high-pileup environment of the HL-LHC?
Key findings
- The convolutional neural network (CNN) achieved the highest significance of 2.85 ± 0.02, significantly outperforming traditional 1D rectangular cuts (0.82 ± 0.02) and other models.
- The particle flow network (PFN) achieved a significance of 1.62 ± 0.01, demonstrating strong performance despite relying on reconstructed jet and particle-level inputs.
- The random forest model achieved a significance of 2.44 ± 0.19, with 544.7 ± 6.3 signal events and 50,000 background events, showing strong signal retention.
- The autoencoder, an unsupervised method, achieved only 0.81 ± 0.01 significance, indicating poor separation between signal and background in the latent space due to overlapping loss distributions.
- The boosted decision tree achieved 1.84 ± 0.09 significance, with 986.3 ± 8.9 signal events and 280,000 background events, showing moderate improvement over classical cuts.
- All results were obtained under zero-pileup conditions, suggesting that reconstruction-dependent models may degrade significantly in high-pileup HL-LHC environments, favoring model-agnostic approaches like CNNs and PFNs.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.