Skip to main content
QUICK REVIEW

[论文解读] The e-value: A Fully Bayesian Significance Measure for Precise Statistical Hypotheses and its Research Program

Julio Michael Stern, Carlos Alberto de Bragança Pereira|arXiv (Cornell University)|Jan 28, 2020
Forecasting Techniques and Applications参考文献 108被引用 4
一句话总结

本文提出了e值,一种用于检验精确(尖锐)假设的完全贝叶斯显著性度量,以及全贝叶斯显著性检验(FBST)框架,该框架严格满足贝叶斯原理的关键要求,如似然性原则和不变性。e值为假设提供了直接、几何且不变的证据度量,在无需人为先验或近似的情况下,能够有效处理零概率假设。

ABSTRACT

This article gives a survey of the e-value, a statistical significance measure a.k.a. the evidence rendered by observational data, X, in support of a statistical hypothesis, H, or, the other way around, the epistemic value of H given X. The $e$-value and the accompanying FBST, the Full Bayesian Significance Test, constitute the core of a research program that was started at IME-USP, is being developed by over 20 researchers worldwide, and has, so far, been referenced by over 200 publications. The e-value and the FBST comply with the best principles of Bayesian inference, including the likelihood principle, complete invariance, asymptotic consistency, etc. Furthermore, they exhibit powerful logic or algebraic properties in situations where one needs to compare or compose distinct hypotheses that can be formulated either in the same or in different statistical models. Moreover, they effortlessly accommodate the case of sharp or precise hypotheses, a situation where alternative methods often require ad hoc and convoluted procedures. Finally, the FBST has outstanding robustness and reliability characteristics, outperforming traditional tests of hypotheses in many practical applications of statistical modeling and operations research.

研究动机与目标

  • 开发一种完全贝叶斯显著性度量,能够严格处理在参数空间零体积子集上定义的尖锐或精确假设。
  • 解决经典p值和其他贝叶斯方法的局限性,这些方法需要人为程序(如为零测度集分配正先验概率)。
  • 提供一种在重新参数化下不变、连续且与似然性原则一致的显著性度量。
  • 通过组合与稳健的程序,实现在复杂模型中的可靠假设检验,特别是在涉及模型比较和结构检验的应用中。
  • 建立一个理论坚实且计算可行的统计推断框架,与基础贝叶斯原则保持一致,并支持医学、运筹学和决策建模等实际应用。

提出的方法

  • e值,记为ev(H|X),定义为在后验分布下,完全包含于原假设H内的最高后验密度(HPD)区域的后验概率。
  • FBST(全贝叶斯显著性检验)使用e值作为检验统计量,基于观测数据X评估支持尖锐假设H的证据。
  • 该方法通过后验分布和在原假设下使后验密度最大化的参数值集合进行几何定义,确保在重新参数化下的不变性。
  • e值的计算不依赖大样本近似,而是基于精确贝叶斯计算,使其成为一种精确程序。
  • 该框架结合主观先验以编码先验知识,并允许使用代数规则在多个模型或假设之间组合证据。
  • 该方法基于一个原则:证据应仅通过似然函数和后验分布进行评估,尊重似然性原则,避免任意阈值或信念比。

实验结果

研究问题

  • RQ1如何构建一种完全贝叶斯显著性度量,使其在无需对零测度集使用非信息或不当先验的情况下,一致评估精确假设?
  • RQ2e值在正则条件下如何满足似然性原则、不变性和连续性,而经典p值或其他贝叶斯检验则不能?
  • RQ3e值如何用于检验涉及多个模型或结构约束的复杂假设,例如贝叶斯网络或影响图中的稀疏性?
  • RQ4在高维或非正则模型中,e值相较于传统频率学派和贝叶斯假设检验具有哪些理论和计算优势?
  • RQ5e值如何推广至半参数和非参数模型,特别是在函数展开或级数近似的情境下?

主要发现

  • e值满足显著性度量的十大关键属性,包括不变性、连续性、遵守似然性原则,以及避免人为假设。
  • e值在原始参数空间中直接提供假设的证据概率度量,直观且可解释,无需变换或近似。
  • FBST在稳健性和可靠性方面优于传统检验,尤其在涉及尖锐假设时,其他方法常失败或需人为调整。
  • e值计算精确,不依赖渐近近似,即使在小样本中也能确保准确性。
  • 该框架支持跨模型的组合与比较,支持对复杂结构假设(如图模型中的稀疏模式)进行检验。
  • 该研究计划已应用于医学诊断、可靠性理论和元分析等多个领域,自提出以来已有超过200篇文献引用该框架。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。