Skip to main content
QUICK REVIEW

[论文解读] Boosted top tagging and its interpretation using Shapley values

Biplob Bhattacherjee, Camellia Bose|arXiv (Cornell University)|Dec 22, 2022
Particle physics theoretical and experimental studies参考文献 47被引用 6
一句话总结

本文提出了一种基于XGBoost的top夸克喷注识别框架,采用N-子喷注张力比和能量相关函数(ECF)可观测量作为输入特征,并通过SHapley Additive exPlanations(SHAP)提升模型可解释性。研究发现,N-子喷注张力特征表现最佳,SHAP分析表明喷注质量与特定D系列项是关键区分特征,且模型在不同top夸克能量下均表现出强鲁棒性。

ABSTRACT

Top tagging has emerged as a fast-evolving subject due to the top quark's significant role in probing physics beyond the standard model. For the reconstruction of top jets, machine learning models have shown a substantial improvement in the classification performance compared to the previous methods. In this work, we build top taggers using $N$-Subjettiness ratios and several Energy Correlation observables as input features to train the eXtreme Gradient BOOSTed decision tree (XGBOOST). The study finds that tighter parton-level matching lead to more accurate tagging. However, in real experimental data, where the parton level data are unknown, this matching cannot be done. We train the XGBOOST models without performing this matching and show that this difference impacts the taggers' effectiveness. Additionally, we test the tagger under different simulation conditions, including changes in center-of-mass energy, parton distribution functions (PDFs), and pileup effects, demonstrating its robustness with performance deviations of less than 1%. Furthermore, we use the SHapley Additive exPlanation (SHAP) framework to calculate the importance of the features of the trained models. It helps us to estimate how much each feature of the data contributed to the model's prediction and what regions are of more importance for each input variable. Finally, we combine all the tagger variables to form a hybrid tagger and interpret the results using the Shapley values.

研究动机与目标

  • 评估基于XGBoost的top夸克喷注识别器在使用N-子喷注张力和能量相关函数(ECF)可观测量作为输入特征时的性能。
  • 研究部分子级top夸克与重建的喷注之间匹配条件(通过ΔRm定义)对喷注识别效率的影响。
  • 评估在不同top夸克横动量(pT)区间下,特别是发生末态辐射导致能量损失时,喷注识别器的鲁棒性。
  • 利用SHAP值解释模型的决策过程,识别最具影响力的特征及其相互作用效应。
  • 构建一个融合所有变量的混合喷注识别器,并通过SHAP分析评估其性能与特征重要性。

提出的方法

  • 使用六组输入特征(N-子喷注张力、C系列、D系列、M系列、N系列、U系列ECF可观测量)训练XGBoost模型。
  • 通过ΔRm(0.6、0.8、1.0)定义匹配条件,评估部分子-喷注匹配对喷注识别性能的影响。
  • 使用SHAP值计算各特征对模型预测的个体贡献,实现对特征重要性的解释。
  • 应用Shapley交互指数量化特征之间的成对交互作用,揭示特征效应如何依赖于共现变量。
  • 通过组合所有输入变量构建混合喷注识别器,并利用SHAP评估其性能与可解释性。
  • 对未修剪与修剪后的喷注进行分析,评估喷注去噪处理对喷注识别性能与特征相关性的影响。

实验结果

研究问题

  • RQ1部分子级top夸克与重建喷注之间匹配半径(ΔRm)的选择如何影响基于XGBoost的top夸克喷注识别器性能?
  • RQ2在基于XGBoost的top夸克喷注识别中,哪组运动学可观测量——N-子喷注张力或ECF——能实现最高的喷注识别效率与AUC?
  • RQ3当初始top夸克横动量发生变化,特别是发生能量损失时,XGBoost喷注识别器对这些变化的鲁棒性如何?
  • RQ4哪些特征对模型分类决策的贡献最为显著,它们之间又如何相互作用?
  • RQ5SHAP值与交互指数如何用于解释并验证基于XGBoost的top夸克喷注识别器的物理驱动决策?

主要发现

  • 在所有六组输入特征中,N-子喷注张力特征在测试准确率与AUC方面均表现最佳,优于基于ECF的特征。
  • 随着ΔRm增大,喷注识别性能下降,未匹配的top喷注表现最差,表明性能对匹配质量高度敏感。
  • 对于修剪后的喷注,基于ECF的特征(尤其是D系列)优于N-子喷注张力特征,尽管整体识别准确率因修剪而下降。
  • XGBoost喷注识别器对初始top夸克能量变化表现出强鲁棒性,在600 GeV top喷注发生能量损失的情况下,性能仅下降0.7%。
  • 在混合喷注识别器中,喷注质量(m_jet)与D3变量的第三项(D_3^{(3.0)})成为最重要的特征,且在多种配置中均排名靠前。
  • SHAP交互分析显示,m_jet与D_3^{(3.0)}之间的交互作用最强,表明二者联合取值对区分top夸克喷注与QCD喷注至关重要。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。