Skip to main content
QUICK REVIEW

[论文解读] The Data-Directed Paradigm for BSM searches

Sergey Volkovich, Federico De Vito Halevy|arXiv (Cornell University)|Jul 27, 2021
Gaussian Processes and Bayesian Inference参考文献 1被引用 7
一句话总结

本文提出了一种数据驱动范式(DDP)用于超出标准模型(BSM)的搜索,该范式利用深度学习在不依赖预定义信号模型的情况下,识别不变质量分布中的显著偏离。通过训练神经网络将数据映射到表示统计显著性的z-分布,该方法在避免耗时的背景和系统不确定性估计的同时,实现了传统轮廓似然比检验95%的灵敏度。

ABSTRACT

In search for new physics, we must leave no stone unturned. We propose a novel data-directed paradigm (DDP) for developing Beyond the Standard Model (BSM) signal hypotheses on the basis of collected data. The DDP is complimentary to traditional theory-directed searches that follow the blind-analysis paradigm and could open the door to regions in the data that will otherwise remain vastly unexplored. We suggest looking directly at the data to identify exclusive selections which exhibit significant deviations from some unique feature of the Standard Model. Such regions should be considered well-motivated for data-directed BSM hypotheses and studied further. In this letter, the paradigm is demonstrated by combining the promising potential of machine learning algorithms with the commonly used search for excesses of events over a smoothly falling background (bump-hunting). A deep neural network is trained to map any invariant mass distribution into a $z$-distribution that indicates the presence of statistically significant bumps. Using this approach, the time and effort consuming tasks of background and systematic uncertainty estimation are avoided. Compared to identifying bumps with the profile-likelihood ratio test using perfectly known background and signal, the algorithm is inferior by only $5$%.

研究动机与目标

  • 开发一种新的BSM搜索范式,以数据驱动取代理论驱动,从而发现标准模型之外的意外新物理。
  • 减少在传统峰迹搜索方法中对耗时的背景建模和系统不确定性估计的依赖。
  • 识别数据分布中可能指示新物理的统计显著偏离,即使这些区域未被传统信号假设覆盖。
  • 证明机器学习可有效替代传统统计检验在峰迹检测中的应用,同时保持高灵敏度。

提出的方法

  • 训练一个深度神经网络,将任意不变质量分布映射为表示与标准模型偏离显著性的z-分布。
  • 网络经过优化,可相对于平滑下降的背景,检测出统计上显著的峰迹——即事件过剩区域。
  • 该方法通过直接从数据中学习显著性,绕过了显式背景建模和系统不确定性估计的需求。
  • z-分布的输出使得无需事先了解信号形状,即可快速识别候选区域以供进一步研究。
  • 通过在已知背景和信号的理想条件下,将该方法与轮廓似然比检验进行性能对比,验证了其有效性。
  • 训练数据被构建为模拟真实的数据分布,包括统计涨落和背景形状。

实验结果

研究问题

  • RQ1使用机器学习的数据驱动方法是否能比传统理论驱动的搜索更有效地检测不变质量分布中的显著偏离?
  • RQ2深度学习在避免背景建模的前提下,能在多大程度上替代传统的统计检验(如轮廓似然比)用于峰迹检测?
  • RQ3当背景和信号完全已知时,所提出方法的灵敏度与轮廓似然比检验相比如何?
  • RQ4该方法能否识别出因模型偏差而被传统搜索遗漏的有希望的BSM信号区域?
  • RQ5在此数据驱动范式中,灵敏度与计算效率之间的权衡如何?

主要发现

  • 当两种方法应用于具有完全已知背景和信号的相同数据时,该数据驱动范式实现了轮廓似然比检验95%的灵敏度。
  • 该方法在无需显式背景建模或系统不确定性估计的情况下,成功识别出不变质量分布中统计显著的峰迹。
  • 深度神经网络能够有效将数据分布映射为反映与标准模型偏离显著性的z-分布。
  • 该方法能够发现传统信号假设未覆盖的数据中此前未探索的区域。
  • 该算法显著减少了在BSM搜索中传统上用于背景估计和不确定性量化的时间与精力。
  • 该方法在发现被理论导向的盲分析方法所忽略的数据区域中的新物理方面,展现出强大潜力。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。