Skip to main content
QUICK REVIEW

[论文解读] Should unfolded histograms be used to test hypotheses?

R. Cousins, S. May|arXiv (Cornell University)|Jul 24, 2016
Gaussian Processes and Bayesian Inference参考文献 4被引用 6
一句话总结

本文通过将展开后的直方图与经过模糊处理的理论预测直接与数据对比(即‘底线测试’),研究了在高能物理中使用展开直方图进行假设检验的可靠性。研究发现,展开过程可能扭曲假设检验结果,尤其是在使用正则化时,因此建议在用于模型比较或拟合优度评估之前,常规性地应用底线测试来验证展开方法的有效性。

ABSTRACT

In many analyses in high energy physics, attempts are made to remove the effects of detector smearing in data by techniques referred to as "unfolding" histograms, thus obtaining estimates of the true values of histogram bin contents. Such unfolded histograms are then compared to theoretical predictions, either to judge the goodness of fit of a theory, or to compare the abilities of two or more theories to describe the data. When doing this, even informally, one is testing hypotheses. However, a more fundamentally sound way to test hypotheses is to smear the theoretical predictions by simulating detector response and then comparing to the data without unfolding; this is also frequently done in high energy physics, particularly in searches for new physics. One can thus ask: to what extent does hypothesis testing after unfolding data materially reproduce the results obtained from testing by smearing theoretical predictions? We argue that this "bottom-line-test" of unfolding methods should be studied more commonly, in addition to common practices of examining variance and bias of estimates of the true contents of histogram bins. We illustrate bottom-line-tests in a simple toy problem with two hypotheses.

研究动机与目标

  • 评估在高能物理中使用展开直方图进行假设检验的有效性。
  • 识别在依赖展开数据进行模型比较或拟合优度检验时可能出现的偏差。
  • 推广使用‘底线测试’——即直接将模糊化的理论预测与数据对比——作为验证展开方法的黄金标准。
  • 强调展开算法中正则化引入的风险,这些风险可能误导假设检验结果。
  • 鼓励在报告中同时提供响应矩阵 $ R $,以便未来可在折叠(模糊)空间中进行比较。

提出的方法

  • 使用响应矩阵 $ R $ 定义展开问题,该矩阵将真实直方图箱内容 $ \vec{\mu} $ 映射到观测(模糊)计数 $ \vec{n} $,其中 $ \vec{\nu} = R\vec{\mu} $。
  • 使用最大似然(ML)和迭代期望最大化(EM)展开算法,从观测数据 $ \vec{n} $ 中估计 $ \hat{\vec{\mu}} $ 及其协方差 $ \hat{U} $。
  • 在展开空间中使用 $ \chi^{2}_{\rm corr} $ 和 $ \Delta\chi^{2}_{\rm corr} $ 等检验统计量进行假设检验,并与模糊空间中的似然比 $ -2\ln\lambda_{0,1} $ 进行比较。
  • 通过模拟理论预测,利用 $ R $ 进行模糊处理,并直接与观测数据 $ \vec{n} $ 对比,不经过展开,从而开展底线测试。
  • 改变关键参数,如高斯模糊宽度 $ \sigma $、真实分布中额外项的振幅 $ B $,以及总事件数,以评估展开结果的稳健性。
  • 将展开空间中的假设检验结果与模糊空间中的结果进行比较,以检测可能表明展开不可靠的差异。

实验结果

研究问题

  • RQ1在展开空间中进行的假设检验是否与将模糊化的理论预测直接与数据对比的检验结果在本质上相同?
  • RQ2展开算法(如 EM 和 ML)中的正则化效应如何影响假设检验结果的可靠性?
  • RQ3探测器分辨率的变化(如 $ \sigma $)或信号振幅($ B $)在多大程度上影响展开结果与底线测试结果之间的一致性?
  • RQ4展开空间中的广义 $ \Delta\chi^{2}_{\rm corr} $ 统计量是否与模糊空间中更强大的 $ -2\ln\lambda_{0,1} $ 统计量等价?
  • RQ5在何种条件下,展开会导致模型比较或拟合优度测试中的误导性结论?

主要发现

  • 在无正则化的矩阵求逆展开中,展开空间中的 $ \Delta\chi^{2}_{\rm corr} $ 统计量与模糊空间中的 $ \chi^{2}_{\rm N} $ 统计量完全相同,而后者已知劣于似然比 $ -2\ln\lambda_{0,1} $。
  • 对于迭代式 EM 和 ML 展开,底线测试揭示了在展开空间与模糊空间中假设检验结果之间存在系统性差异,表明基于展开结果的推断可能存在不可靠性。
  • 真实分布中额外项的振幅 $ B $ 显著影响展开结果与底线测试结果之间的一致性,尤其是在 EM 展开中更为明显。
  • 直方图中的事件数量影响展开假设检验的可靠性,统计量越低,检验结果的差异越明显。
  • EM 展开中的迭代次数影响底线测试结果,收敛行为会影响最终检验统计量的有效性。
  • 本研究发现,展开算法中正则化引入的偏差可能严重扭曲假设检验结果,这是一个关键问题,需在依赖展开数据进行模型比较前进一步研究。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。