[论文解读] Quark jet versus gluon jet: fully-connected neural networks with high-level features
该论文提出使用全连接神经网络(FNNs)结合专家设计的高阶喷注可观测量来分类夸克与胶子喷注,在粒子层次数据上性能与卷积神经网络(DCNNs)相当,在探测器层次数据上仅出现约15%的性能下降。在宽 $p_{TJ}$ 范围($[200, 1000]$ GeV)上训练的FNN与在单个 $p_{TJ}$ 分箱上训练的性能几乎完全一致,且仅需14个可观测量的最小集合即可捕获几乎全部判别能力。
Jet identification is one of the fields in high energy physics that machine learning has begun to make an impact. More often than not, convolutional neural networks are used to classify jet images with the benefit that essentially no physics input is required. Inspired by a recent work by Datta and Larkoski, we study the classification of quark/gluon-initiated jets based on fully-connected neural networks (FNNs), where expert-designed physical variables are taken as input. FNNs are applied in two ways: trained separately on various narrow jet transverse momentum $p_{TJ}$ bins; trained on a wide region of $p_{TJ} \in [200,~1000]$ GeV. We find their performances are almost the same. The performance is better when the $p_{TJ}$ is larger. Jet discrimination with FNN is studied on both particle and detector level data. The results based on particle level data are comparable with those from deep convolutional neural networks, while the significance improvement characteristic (SIC) from detector level data would at most decrease by $15\%$. We also test the performance of FNNs with full set or subsets of jet observables as input features. The FNN with one subset consisting of fourteen observables shows nearly no degradation of performance. This indicates that these fourteen expert-designed observables could have captured the most necessary information for separating quark and gluon jets.
研究动机与目标
- 评估全连接神经网络(FNNs)在使用物理启发的高阶可观测量进行夸克与胶子喷注判别中的性能。
- 比较FNN在不同喷注横动量($p_{TJ}$)区域的性能表现,包括在宽 $p_{TJ}$ 范围上训练与在单独分箱上训练的对比。
- 评估使用快速模拟(Delphes)时探测器效应对基于FNN的喷注分类性能的影响。
- 确定是否存在一个最小可观测量子集,可保留接近最优的分类性能。
- 探究浅层FNN是否能在该喷注标记任务中达到深层网络的性能水平。
提出的方法
- 使用六层隐藏层、每层300个神经元的FNN,在粒子层次和探测器层次的喷注数据上进行训练,使用36个专家设计的喷注可观测量,包括喷注质量、N-子喷注性、能量相关函数和粒子多重性。
- 在窄 $p_{TJ}$ 分箱($[200,220]$、$[500,550]$、$[1000,1100]$ GeV)上分别训练FNN,并训练一个在宽 $p_{TJ}$ 范围($[200,1000]$ GeV)上训练的FNN,以比较在不同动量尺度下的泛化能力。
- 利用Delphes模拟探测器效应,以估计夸克-胶子分离在统计显著性(SIC)上的性能下降。
- 通过ROC AUC和在50%夸克喷注接受率下的胶子喷注效率评估性能,并对可观测量子集进行消融实验,以识别最小有效集合。
- 将单隐藏层(300个神经元)的浅层FNN与深层FNN进行比较,以评估该分类任务中深度的必要性。
- 研究利用 $\tau_{1}^{(0.5)}$、$\tau_{2}^{(1)}$、$U_{1}^{(0.5)}$、$\lambda_{1}^{(0.5)}$ 和粒子多重性等可观测量,以编码喷注子结构和色荷差异。
实验结果
研究问题
- RQ1在高阶、专家设计的喷注可观测量上训练的全连接神经网络能否实现与卷积神经网络相当的夸克-胶子喷注判别性能?
- RQ2在宽 $p_{TJ}$ 范围($[200,1000]$ GeV)上训练单一FNN是否能获得与在窄 $p_{TJ}$ 分箱上分别训练FNN相当的性能?
- RQ3在使用FNN进行分类时,探测器分辨率与接受度的退化会使夸克-胶子分离的统计显著性降低多少?
- RQ4是否存在一个最小可观测量子集,可保留接近最大化的分类性能?
- RQ5单隐藏层的浅层FNN是否能与深层FNN在夸克-胶子喷注标记任务中达到相同性能?
主要发现
- 对于 $p_{TJ} \in [1000,1100]$ GeV,FNN的ROC AUC达到0.899,优于同一 $p_{TJ}$ 分箱中DCNN的性能。
- 在 $p_{TJ} \in [1000,1100]$ GeV下,50%夸克喷注接受率时的胶子喷注效率为2.8%,表明在高 $p_{TJ}$ 下具有强大的判别能力。
- 在宽 $p_{TJ}$ 范围($[200,1000]$ GeV)上训练的FNN性能与在单个 $p_{TJ}$ 分箱上训练的几乎完全一致,表明具有良好的泛化能力。
- 探测器效应导致统计显著性提升特征(SIC)最多下降15%,表明对真实探测器分辨率与接受度具有鲁棒性。
- 仅包含14个可观测量的最小集合——包括喷注质量、$\tau_{1}^{(0.5)}$、$\tau_{2}^{(1)}$、$\tau_{3}^{(2)}$、$U_{1}^{(0.5)}$、$\lambda_{1}^{(0.5)}$ 和多重性——的性能与完整36个可观测量集合几乎完全一致。
- 单隐藏层(300个神经元)的浅层FNN性能与六层深层FNN相当,表明所选可观测量已包含足够判别信息,使简单模型即可达到最优性能。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。