Skip to main content
QUICK REVIEW

[Paper Review] Bidirectional Long Short-Term Memory (BLSTM) neural networks for reconstruction of top-quark pair decay kinematics

Fardin Syed, R. Di Sipio|arXiv (Cornell University)|Sep 3, 2019
Particle physics theoretical and experimental studies4 citations
TL;DR

This paper proposes a bidirectional Long Short-Term Memory (BLSTM) neural network, named AngryTops, to reconstruct top-quark pair decay kinematics in the muon+jets channel at 13 TeV proton-proton collisions. Trained on Monte Carlo events simulated with MadGraph5, Pythia8, and Delphes3, the model achieves reconstruction performance comparable to the standard $χ^2$-fit method, demonstrating the viability of deep learning for probabilistic kinematic reconstruction in high-energy physics.

ABSTRACT

A probabilistic reconstruction using machine-learning of the decay kinematics of top-quark pairs produced in high-energy proton-proton collisions is presented. A deep neural network whose core consists of a Bidirectional Long Short-Term Memory (BLSTM) is trained to infer the four-momenta of the two top quarks produced in the hard scattering process. The MadGraph5+Pythia8 Monte Carlo event generator is used to create a sample of top-quark pairs decaying in the $\\mu$+jets channel, whose final-state objects are used to create the input to the deep neural network. Distortions due to limited resolution of the experimental apparatus are simulated with the Delphes3 fast detector simulator. The level of agreement between the Monte Carlo predictions and the BLSTM for kinematic distributions at parton level is comparable to that obtained using a benchmark method that finds the jet permutation that minimizes an objective function.

Motivation & Objective

  • To develop a machine learning-based approach for reconstructing top-quark pair decay kinematics in high-energy proton-proton collisions.
  • To overcome limitations of traditional reconstruction methods like $χ^2$-fit, which fail when jets fall below transverse momentum thresholds due to detector resolution.
  • To evaluate whether a BLSTM-based neural network can outperform or match standard algorithms in reconstructing four-momenta of top quarks, W bosons, and bottom quarks.
  • To investigate the robustness and generalization of the model across different kinematic distributions, especially in regions with low jet $p_{\rm T}$.
  • To enable flexible, extrapolative reconstruction that does not rely on parametric transfer functions or fixed jet permutations.

Proposed method

  • A deep neural network with a Bidirectional Long Short-Term Memory (BLSTM) core is trained to infer the four-momenta of top quarks from final-state objects in the $μ$+jets $t\bar{t}$ decay channel.
  • Monte Carlo events are generated using MadGraph5 for matrix element calculations and Pythia8 for parton showering, with MLM matching to ensure consistency.
  • Detector effects are simulated using Delphes3, including pile-up from additional $pp$ interactions, to reproduce realistic LHC conditions.
  • Input features include reconstructed muons, jets, and missing transverse energy, with $p_{\rm T} > 20$ GeV and $|\eta| < 2.5$ cuts applied to key particles.
  • The network is trained to predict the four-momenta of the two top quarks, their decay products (W bosons and bottom quarks), and the corresponding kinematic distributions.
  • Performance is evaluated by comparing the reconstructed distributions to parton-level Monte Carlo truth, using metrics such as agreement in $p_{\rm T}$, $\eta$, and $\phi$ distributions.

Experimental results

Research questions

  • RQ1Can a BLSTM-based neural network reconstruct top-quark pair decay kinematics with performance comparable to the standard $χ^2$-fit method?
  • RQ2How does the model perform in reconstructing kinematic variables when jets fall below the $p_{\rm T} > 20$ GeV threshold, where traditional methods fail?
  • RQ3To what extent does the BLSTM model learn the $p_{\rm T}$ cutoff and shape of transverse momentum distributions compared to Monte Carlo truth?
  • RQ4What are the limitations of the BLSTM approach in reconstructing $W$-boson and $b$-quark kinematics, and how do they compare to those of the $χ^2$-fit?
  • RQ5How do angular distributions ($\eta$, $\phi$) reconstructed by the neural network compare to Monte Carlo truth, and what causes observed asymmetries?

Key findings

  • The BLSTM-based model, AngryTops, achieves reconstruction performance for top-quark $p_{\rm T}$ and $\eta$ distributions that is comparable to the benchmark $χ^2$-fit method.
  • Both AngryTops and $χ^2$-fit under-estimate the top-quark $p_{\rm T}$ in a similar manner, suggesting a shared systematic bias in the reconstruction process.
  • AngryTops consistently underestimates the transverse momentum of $b$-quarks and fails to learn the $p_{\rm T} > 20$ GeV threshold, even though all training jets meet this criterion.
  • The model shows asymmetries in $\phi$ and rapidity distributions that are not present in the $χ^2$-fit, and these asymmetries vary between training sessions, indicating potential instability or incomplete training.
  • While the $W$-boson and $b$-quark kinematics are poorly reconstructed by both methods, the neural network does not improve on the $χ^2$-fit in these observables.
  • The $χ^2$-fit slightly overestimates $p_{\rm T}$ and correctly accounts for the $p_{\rm T}$ cutoff in $b$-quark distributions, outperforming AngryTops in $\phi$ distribution accuracy.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.