Skip to main content
QUICK REVIEW

[Paper Review] Uncovering hidden patterns in collider events with Bayesian probabilistic models

Darius A. Faroughy|arXiv (Cornell University)|Dec 15, 2020
Topic Modeling7 references4 citations
TL;DR

This paper proposes a Bayesian probabilistic model based on Latent Dirichlet Allocation (LDA) to uncover hidden new physics patterns in collider events without labeled data. By treating jet substructure measurements as exchangeable, discrete tokens in the Lund plane, the model learns two latent 'themes'—one capturing QCD background and the other identifying a $W'\!-\!\phi$ BSM signal at 5% signal-to-background ratio, demonstrating successful unsupervised discovery of new physics features.

ABSTRACT

Individual events at high-energy colliders like the LHC can be represented by a sequence of measurements, or 'point patterns' in an observable space. Starting from this data representation, we build a simple Bayesian probabilistic model for event measurements useful for unsupervised event classification in beyond the standard model (BSM) studies. In order to arrive to this model we assume that the event measurements are exchangeable (and apply De Finetti's representation theorem), the data is discrete, and measurements are generated from multiple 'latent' distributions, called 'themes'. The resulting probabilistic model for collider events is a mixed-membership model known as Latent Dirichlet Allocation (LDA), a model extensively used in natural language processing applications. By training on point patterns in the primary Lund plane, we demonstrate that a two-theme LDA model can learn to distinguish in unlabelled dijet events the hidden new physics patterns produced by a BSM signature from a much larger QCD background. Based on 1904.04200 and 2005.12319.

Motivation & Objective

  • To develop a simple, interpretable Bayesian probabilistic model for unsupervised event classification in high-energy collider physics.
  • To detect hidden new physics signatures in jet substructure data without prior labeling.
  • To apply exchangeability and discrete tokenization to model collider events as a mixed-membership process.
  • To evaluate whether LDA can extract rare BSM signals from dominant QCD backgrounds in unlabelled event samples.
  • To demonstrate the feasibility of using natural language processing-inspired models for high-energy physics event analysis.

Proposed method

  • Represent collider events as exchangeable sequences of measurements in observable space, enabling use of De Finetti’s representation theorem.
  • Discretize continuous observables (e.g., Lund plane variables) into bins to model measurements as discrete tokens.
  • Apply Latent Dirichlet Allocation (LDA), a mixed-membership model, to learn latent 'themes' representing underlying probability distributions over the binned observable space.
  • Use a Dirichlet prior with asymmetric concentration parameters ($\alpha_0 \approx 0.9$, $\alpha_2 \approx 0.1$) to bias one theme toward QCD-like patterns and the other toward rare BSM features.
  • Train the LDA model on 100k unlabelled dijet events with a $W'\!-\!\phi$ signal embedded at 5% signal-to-background ratio using the Gensim library.
  • Use the learned themes to reconstruct and compare with truth-level distributions for signal and background in the primary Lund plane.

Experimental results

Research questions

  • RQ1Can exchangeability and discrete tokenization of collider event measurements enable a principled Bayesian model for unsupervised event classification?
  • RQ2Can LDA effectively learn and separate QCD background patterns from rare BSM signal features in unlabelled dijet events?
  • RQ3Does a two-theme LDA model trained on a mixed sample with only 5% signal-to-background ratio successfully recover the true signal structure in the Lund plane?
  • RQ4Can the latent themes in LDA be interpreted as physically meaningful distributions corresponding to known QCD and BSM jet substructure patterns?
  • RQ5To what extent can a model inspired by natural language processing detect subtle, non-uniform features in sparse, irregularly shaped collider event patterns?

Key findings

  • The two-theme LDA model successfully learned a first theme that closely matches the truth-level QCD background distribution in the primary Lund plane.
  • The second theme captured the distinctive two-cluster structure in the Lund plane associated with the $W'\!\to\!W\phi\to\!WW\!\to\!jjjj$ decay chain, characteristic of the BSM signal.
  • Despite a signal-to-background ratio of only 5%, the model identified and isolated the non-QCD signal pattern in an unsupervised manner.
  • The learned themes demonstrated that LDA can extract physically relevant features from sparse, irregularly shaped event patterns in the Lund plane.
  • The model’s performance was validated by visual comparison of learned themes with truth-level distributions, showing strong alignment with both background and signal structures.
  • The results confirm that LDA, when applied to discretized jet substructure observables, enables effective unsupervised discovery of new physics signatures in high-energy collider data.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.