Skip to main content
QUICK REVIEW

[Paper Review] Learning the EFT likelihood with tree boosting

S. Chatterjee, Stefan Rohshap|arXiv (Cornell University)|May 25, 2022
Nuclear reactor physics and engineering4 citations
TL;DR

This paper introduces a tree-boosting algorithm, Boosted Information Trees (BIT), to learn the likelihood ratio in effective field theory (EFT) analyses for multi-parameter measurements at the LHC. By training on simulated event data with per-event information, BIT approximates the differential cross section ratio order-by-order in Wilson coefficients, achieving near-optimal statistical power comparable to the likelihood ratio test, with significant improvements over baseline $p_{\textrm{T,Z}}$-based methods in exclusion reach.

ABSTRACT

We develop a tree boosting algorithm for collider measurements of multiple Wilson coefficients in effective field theories describing phenomena beyond the standard model of particle physics. The design of the discriminant exploits per-event information of the simulated data sets that encodes the predictions for different values of the Wilson coefficients. This ``Boosted Information Tree'' algorithm provides nearly optimal discrimination power order-by-order in the expansion in the Wilson coefficients and approaches the optimal likelihood ratio test statistic. As a proof-of-principle, we apply the algorithm to the $ extrm{pp} ightarrow extrm{Zh}$ process for different types of modeling.

Motivation & Objective

  • To develop a statistically optimal, scalable method for measuring multiple Wilson coefficients in SM-EFT using collider data.
  • To overcome the challenge of high-dimensional feature spaces and complex parameter spaces in global EFT fits.
  • To enable efficient, order-by-order estimation of the differential cross section ratio using tree-based regression.
  • To demonstrate improved exclusion sensitivity compared to conventional $p_{\textrm{T,Z}}$-based strategies in realistic simulation settings.
  • To establish a framework that leverages per-event simulation data to train models that approximate the true detector-level likelihood.

Proposed method

  • The method uses tree boosting to regress on the differential cross section ratio $\hat{R}(\boldsymbol{x}|\boldsymbol{\theta}, \boldsymbol{\theta}_0)$ as a function of event features $\boldsymbol{x}$, using a mean-squared error (MSE) loss functional.
  • It expands the cross section ratio in a Taylor series around $\boldsymbol{\theta}_0$, with each order estimated by a separate boosted tree regressor.
  • The weak learner is a CART decision tree, trained on simulated events weighted by the SM-EFT dependence to encode the coefficient functions.
  • The boosting algorithm sequentially improves the estimate of $\hat{R}(\boldsymbol{x}|\boldsymbol{\theta}, \boldsymbol{\theta}_0)$ by minimizing the MSE loss over the training data.
  • The final test statistic $\hat{q}(\mathcal{D}) = -\log \hat{R}(\boldsymbol{x}|\boldsymbol{\theta}, \boldsymbol{\theta}_0)$ is used for hypothesis testing and exclusion limits.
  • A binned version of the estimator is introduced to maintain optimality even when higher-order EFT terms are uncertain, by decoupling training from the coefficients of interest.

Experimental results

Research questions

  • RQ1Can tree boosting be used to learn the likelihood ratio in multi-parameter EFT measurements with near-optimal statistical power?
  • RQ2How well does the Boosted Information Tree (BIT) method approximate the true likelihood ratio under realistic LHC simulation conditions?
  • RQ3Does BIT outperform conventional $p_{\textrm{T,Z}}$-based analysis strategies in terms of exclusion reach for Wilson coefficients?
  • RQ4To what extent does the method remain optimal when higher-order terms in the Wilson coefficient expansion are neglected or uncertain?
  • RQ5Can the BIT framework be applied to realistic event generators including backgrounds and detector effects?

Key findings

  • The BIT method achieves nearly optimal discrimination power in hypothesis testing, approaching the performance of the true likelihood ratio test statistic under the Neyman-Pearson lemma.
  • In a toy model of the $\mathrm{pp} \to \mathrm{Z}h$ process, the BIT-based test statistic shows optimal power and accurate exclusion contours, validating its theoretical optimality.
  • In a realistic MadGraph5_aMC@NLO simulation including backgrounds, BIT significantly improves exclusion reach compared to a $p_{\textrm{T,Z}}$-based strategy with a neural network classifier.
  • The binned version of the BIT estimator maintains optimal performance even when higher-order EFT terms are uncertain, by decoupling training from the coefficients of interest.
  • The method provides a simple, analytic test statistic with direct interpretability, enabling both binned and unbinned hypothesis testing.
  • The algorithm scales efficiently with the number of Wilson coefficients, requiring only as many regressors as the order of the EFT expansion.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.