Skip to main content
QUICK REVIEW

[论文解读] Model-Independent Detection of New Physics Signals Using Interpretable Semi-Supervised Classifier Tests

Purvasha Chakravarti, Mikael Kuusela|arXiv (Cornell University)|Feb 15, 2021
Gaussian Processes and Bayesian Inference被引用 6
一句话总结

该论文提出了一种模型无关的、半监督分类器框架,用于在不假设特定信号模型的前提下,检测高维粒子物理数据中的新物理信号。通过使用基于分类器的检验统计量——似然比、AUC 和误分类误差,该方法在信号出乎意料或建模错误时,相较于模型依赖方法,对标准模型背景的偏离检测能力更强,这一优势在希格斯玻色子搜索模拟中得到验证。

ABSTRACT

A central goal in experimental high energy physics is to detect new physics signals that are not explained by known physics. In this paper, we aim to search for new signals that appear as deviations from known Standard Model physics in high-dimensional particle physics data. To do this, we determine whether there is any statistically significant difference between the distribution of Standard Model background samples and the distribution of the experimental observations, which are a mixture of the background and a potential new signal. Traditionally, one also assumes access to a sample from a model for the hypothesized signal distribution. Here we instead investigate a model-independent method that does not make any assumptions about the signal and uses a semi-supervised classifier to detect the presence of the signal in the experimental data. We construct three test statistics using the classifier: an estimated likelihood ratio test (LRT) statistic, a test based on the area under the ROC curve (AUC), and a test based on the misclassification error (MCE). Additionally, we propose a method for estimating the signal strength parameter and explore active subspace methods to interpret the proposed semi-supervised classifier in order to understand the properties of the detected signal. We also propose a Score test statistic that can be used in the model-dependent setting. We investigate the performance of the methods on a simulated data set related to the search for the Higgs boson at the Large Hadron Collider at CERN. We demonstrate that the semi-supervised tests have power competitive with the classical supervised methods for a well-specified signal, but much higher power for an unexpected signal which might be entirely missed by the supervised tests.

研究动机与目标

  • 解决高能物理中模型依赖搜索方法在无法检测意外或建模错误的新物理信号时的局限性。
  • 开发一种模型无关的方法,仅使用背景数据和实验数据检测新物理信号,无需依赖模拟的信号样本。
  • 提升对新颖或未预期信号的检测能力,这些信号常因模型误设而被传统监督分类器遗漏。
  • 通过主动子空间方法和信号强度估计,实现对信号的可解释表征。
  • 在信号模型不确定或未知的情况下,为经典似然比检验提供一种稳健的替代方案。

提出的方法

  • 构建三种基于半监督分类器的检验统计量:估计的似然比检验(LRT)、ROC曲线下面积(AUC)和误分类误差(MCE),用于高维两样本检验。
  • 使用在背景数据和未标记实验数据上训练的半监督分类器检测偏离,无需信号模拟。
  • 应用主动子空间方法解释分类器,识别驱动信号检测的关键数据维度。
  • 提出一种新方法估计信号强度参数,量化实验样本中信号事件的比例。
  • 为模型依赖场景提出一种Score检验统计量,以与经典方法进行比较。
  • 利用LHC提供的16维模拟希格斯玻色子搜索数据集,在真实条件下评估性能。

实验结果

研究问题

  • RQ1当信号模型未知或错误指定时,模型无关的分类器检验能否检测到新物理信号?
  • RQ2在信号模型错误指定的情况下,半监督分类器检验的检验力与经典模型依赖似然比检验相比如何?
  • RQ3主动子空间方法在多大程度上能够解释高维数据中导致信号检测的关键特征?
  • RQ4在缺乏信号模拟数据的情况下,能否可靠地进行信号强度估计?
  • RQ5所提出的方法在信号模型正确指定时是否保持竞争性检测能力,同时在意外信号上显著优于模型依赖方法?

主要发现

  • 当信号模型正确指定时,所提出的半监督分类器检验在检测能力上与经典模型依赖方法相当。
  • 当信号模型错误指定时,模型依赖方法完全无法检测到信号,而模型无关方法仍保持高检测能力。
  • 主动子空间方法成功识别出对信号检测贡献最大的数据维度,实现了可解释的信号表征。
  • 即使无信号模拟数据,信号强度估计仍可行且准确,提供了信号存在的定量度量。
  • 基于AUC和MCE的检验统计量在多种信号配置下表现出稳健性和高检测力,尤其在高维设置中表现突出。
  • 该方法可在不了解新粒子属性的前提下实现新粒子的探测,为探索性物理搜索提供了关键优势。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。