[论文解读] Changing the paradigm of fixed significance levels: Testing Hypothesis by Minimizing Sum of Errors Type I and Type II
本文通过用最小化第一类与第二类错误加权和的方法取代固定的显著性水平,实现了假设检验范式的根本性转变。该方法基于预测匹配先验的贝叶斯因子,实现了最优且一致的检验,尊重似然原理,并调和了频率学派与贝叶斯学派的观点,为p值和固定α水平提供了一种更平衡、更符合科学逻辑的替代方案。
Our purpose, is to put forward a change in the paradigm of testing by generalizing a very natural idea exposed by Morris DeGroot (1975) aiming to an approach that is attractive to all schools of statistics, in a procedure better suited for the needs of science. DeGroot's seminal idea is to base testing statistical hypothesis on minimizing the weighted sum of type I plus type II error instead of of the prevailing paradigm which is fixing type I error and minimizing type II error. DeGroot's result is that in simple vs simple hypothesis the optimal criterion is to reject, according to the likelihood ratio as the evidence (ordering) statistics using a fixed threshold value, instead of a fixed tail probability. By defining expected type I and type II errors, we generalize DeGroot's approach and find that the optimal region is defined by the ratio of evidences, that is, averaged likelihoods (with respect to a prior measure) and a threshold fixed. This approach yields an optimal theory in complete generality, which the Classical Theory of Testing does not. This can be seen as a Bayes-Non-Bayes compromise: the criteria (weighted sum of type I and type II errors) is Frequentist, but the test criterion is the ratio of marginalized likelihood, which is Bayesian. We give arguments, to push the theory still further, so that the weighting measures (priors)of the likelihoods does not have to be proper and highly informative, but just predictively matched, that is that predictively matched priors, give rise to the same evidence (marginal likelihoods) using minimal (smallest) training samples. The theory that emerges, similar to the theories based on Objective Bayes approaches, is a powerful response to criticisms of the prevailing approach of hypothesis testing, see for example Ioannidis (2005) and Siegfried (2010) among many others.
研究动机与目标
- 解决传统假设检验中存在的科学不一致与失衡问题,即固定的第一类错误率导致第二类错误失控,尤其在大样本中更为显著。
- 将Morris DeGroot于1975年提出的最小化第一类与第二类错误加权和的思想,从简单对简单假设推广至复合模型。
- 开发一种通用的、最优的检验框架,该框架与似然原理和可选停止原则一致,调和频率学派与贝叶斯学派的方法。
- 证明即使使用不当先验,只要其与最小训练样本预测匹配,即可通过边际似然获得有效且一致的证据,而无需使用 proper 先验。
- 通过展示一种方法可避免p值和固定α水平所引发的广泛批评,特别是当p值极小但实际意义可忽略时仍拒绝原假设的悖论。
提出的方法
- 通过在参数空间上定义先验分布来计算期望错误,从而最小化第一类与第二类错误的加权和。
- 最优拒绝域由边际似然(证据)之比确定,该比值在使用 proper 先验时即对应贝叶斯因子。
- 采用在每个假设下均为常数的损失函数,从而得出贝叶斯因子作为最优检验统计量。
- 为考虑实际显著性,引入无感区域,即差异低于阈值Δ时被视为可忽略,相应调整证据比。
- 使用预测匹配的不当先验——即基于最小训练样本校准的不当先验——以确保一致性与有效性,而无需依赖 proper 先验。
- 理论基础源于决策理论,检验准则为将平均似然(证据)之比与固定阈值比较,而非固定p值或α水平。
实验结果
研究问题
- RQ1能否开发一种通用的、最优的假设检验框架,在不预先固定α的情况下平衡第一类与第二类错误?
- RQ2与传统的Neyman-Pearson检验相比,最小化错误加权和的方法在一致性和错误控制方面表现如何?
- RQ3当不当先验通过最小训练数据进行预测匹配时,其在假设检验中的意义有多大?
- RQ4为何p值在大样本中会导致误导性结论?如何通过基于证据的标准避免此类问题?
- RQ5能否使一种检验既最优又一致,同时尊重似然原理与可选停止原则?
主要发现
- 在样本量N=104,490,000、成功次数S=52,263,471的大样本二项分布示例中,p值为0.0003,导致拒绝原假设,但使用Jeffreys先验的贝叶斯因子为18.7,表明对原假设存在强烈支持。
- 当实际显著性阈值Δ=0.0005(0.05%)时,对备择假设的证据比达到5.1×10^10,表明尽管p值极小,原假设仍获得压倒性支持。
- 该方法确保一致性:当证据增加时,若最小化加权和,则第一类与第二类错误均收敛于零。
- 最优检验满足似然原理与可选停止原理,而p值方法则不满足。
- 在最优检验下,第一类与第二类错误之比被限制在b/a以内,表明该方法内在地控制了相对错误权衡。
- 即使使用不当先验,只要其为预测匹配,所得证据(边际似然)仍有效,并导出相同的最优决策规则。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。