Skip to main content
QUICK REVIEW

[論文レビュー] Model-Independent Detection of New Physics Signals Using Interpretable Semi-Supervised Classifier Tests

Purvasha Chakravarti, Mikael Kuusela|arXiv (Cornell University)|Feb 15, 2021
Gaussian Processes and Bayesian Inference被引用数 6
ひとこと要約

本論文は、特定の信号モデルを仮定せずに、高次元の素粒子物理学データにおける新しい物理現象の信号を検出するモデルに依存しない半教師あり分類器フレームワークを提案する。分類器に基づく検定統計量(尤度比、AUC、誤分類誤差)を用いることで、信号が予期しないか誤って指定された場合に、モデルに依存する手法よりも高い検出力を持つ。これはヒッグス粒子探索のシミュレーションで実証された。

ABSTRACT

A central goal in experimental high energy physics is to detect new physics signals that are not explained by known physics. In this paper, we aim to search for new signals that appear as deviations from known Standard Model physics in high-dimensional particle physics data. To do this, we determine whether there is any statistically significant difference between the distribution of Standard Model background samples and the distribution of the experimental observations, which are a mixture of the background and a potential new signal. Traditionally, one also assumes access to a sample from a model for the hypothesized signal distribution. Here we instead investigate a model-independent method that does not make any assumptions about the signal and uses a semi-supervised classifier to detect the presence of the signal in the experimental data. We construct three test statistics using the classifier: an estimated likelihood ratio test (LRT) statistic, a test based on the area under the ROC curve (AUC), and a test based on the misclassification error (MCE). Additionally, we propose a method for estimating the signal strength parameter and explore active subspace methods to interpret the proposed semi-supervised classifier in order to understand the properties of the detected signal. We also propose a Score test statistic that can be used in the model-dependent setting. We investigate the performance of the methods on a simulated data set related to the search for the Higgs boson at the Large Hadron Collider at CERN. We demonstrate that the semi-supervised tests have power competitive with the classical supervised methods for a well-specified signal, but much higher power for an unexpected signal which might be entirely missed by the supervised tests.

研究の動機と目的

  • 予期しないか誤って指定された新物理信号を検出できない、高エネルギー物理学におけるモデルに依存する探索の限界を克服すること。
  • シミュレートされた信号サンプルに依存せずに、バックグラウンドデータと実験データのみを用いて新物理信号を検出するモデルに依存しない手法を開発すること。
  • 従来の教師あり分類器がモデルの誤指定により見逃す、新しいまたは予期しない信号の検出力を向上させること。
  • アクティブサブスパイス法と信号強度推定を用いて、解釈可能な信号の特徴を明らかにすること。
  • 信号モデルが不確実または未知の状況において、古典的な尤度比検定の代替手段を提供すること。

提案手法

  • 背景データとラベルなしの実験データを用いて訓練された半教師あり分類器から、推定尤度比検定(LRT)、受信者操作特性曲線下の面積(AUC)、誤分類誤差(MCE)の3つの検定統計量を構築し、高次元の2標本検定に応用する。
  • 信号のシミュレーションに依存せず、バックグラウンドデータとラベルなしの実験データから構築された半教師あり分類器を用いて、信号の存在を検出する。
  • アクティブサブスパイス法を適用し、信号検出に寄与する最も重要なデータ次元を特定する。
  • 実験サンプルに含まれる信号イベントの割合を定量化する新しい手法を用いて、信号強度パラメータを推定する。
  • 古典的手法との比較を可能にするために、モデルに依存する設定におけるスコア検定統計量を提案する。
  • LHCの16次元のヒッグス粒子探索のシミュレーションデータセットを用い、現実的な条件下での性能を評価する。

実験結果

リサーチクエスチョン

  • RQ1信号モデルが不明または誤って指定された場合に、モデルに依存しない分類器検定は新物理信号を検出できるか?
  • RQ2信号モデルが誤って指定された場合に、半教師あり分類器検定の検出力は、古典的なモデルに依存する尤度比検定と比べてどの程度優れているか?
  • RQ3アクティブサブスパイス法は、高次元データにおける信号検出に寄与する特徴をどの程度解釈できるか?
  • RQ4信号シミュレーションデータが利用できない状況でも、信号強度推定は信頼性を持って実行可能か?
  • RQ5信号モデルが正しく指定された場合でも、提案手法は競争力のある検出力を維持するか。また、予期しない信号に対しては、モデルに依存する手法を著しく上回るか?

主な発見

  • 信号モデルが正しく指定された場合、提案された半教師あり分類器検定は古典的なモデルに依存する手法と同等の検出力を達成する。
  • 信号モデルが誤って指定された場合、モデルに依存する手法は信号を完全に検出できなくなるが、モデルに依存しないアプローチは高い検出力を維持する。
  • アクティブサブスパイス法は、信号検出に寄与する最も関連性の高いデータ次元を的確に特定でき、解釈可能な信号特徴の特定を可能にする。
  • 信号シミュレーションデータが入手不可であっても、信号強度推定は実行可能で、高い正確性を示す。これにより、信号の存在の定量的測定が可能になる。
  • AUCに基づくおよびMCEに基づく検定統計量は、多様な信号設定、特に高次元設定において、頑健で高い検出力を示す。
  • 本手法により、粒子の性質を事前に知らなくても新粒子の検出が可能となり、探索的物理学探索における重要な利点を提供する。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。