[论文解读] Statistical inference optimized with respect to the observed sample for single or multiple comparisons
本文提出了一种基于归一化最大似然(NML)比值的新颖统计证据度量——区分信息(DI),用于量化对简单原假设的支持或反对证据,且无需依赖先验分布。该方法具有极小化最大最优性,具有渐近可解释性,并对偶然信息具有鲁棒性,即使在小样本量下对多重比较的敏感性也极低,从而减少了对传统p值校正的需求。
The normalized maximum likelihood (NML) is a recent penalized likelihood that has properties that justify defining the amount of discrimination information (DI) in the data supporting an alternative hypothesis over a null hypothesis as the logarithm of an NML ratio, namely, the alternative hypothesis NML divided by the null hypothesis NML. The resulting DI, like the Bayes factor but unlike the p-value, measures the strength of evidence for an alternative hypothesis over a null hypothesis such that the probability of misleading evidence vanishes asymptotically under weak regularity conditions and such that evidence can support a simple null hypothesis. Unlike the Bayes factor, the DI does not require a prior distribution and is minimax optimal in a sense that does not involve averaging over outcomes that did not occur. Replacing a (possibly pseudo-) likelihood function with its weighted counterpart extends the scope of the DI to models for which the unweighted NML is undefined. The likelihood weights leverage side information, either in data associated with comparisons other than the comparison at hand or in the parameter value of a simple null hypothesis. Two case studies, one involving multiple populations and the other involving multiple biological features, indicate that the DI is robust to the type of side information used when that information is assigned the weight of a single observation. Such robustness suggests that very little adjustment for multiple comparisons is warranted if the sample size is at least moderate.
研究动机与目标
- 开发一种可解释的统计证据度量,能够支持简单原假设,克服p值和贝叶斯因子的局限性。
- 通过引入极小化最大最优、无需先验的方法,解决多重比较中误导性证据的问题。
- 确保在使用其他比较的偶然信息时,该度量仍保持鲁棒性,即使样本量较小。
- 在不同数量的比较下提供一致的证据解释,不同于依赖于检验次数的p值或后验概率。
- 为科学家提供一种可靠、客观的工具,以评估数据是否强烈支持备择假设、原假设或两者皆不支持。
提出的方法
- 区分信息(DI)定义为备择假设与原假设下归一化最大似然(NML)值比值的对数。
- NML惩罚似然确保了极小化最大最优性,并避免对未观测结果的平均,使该方法具有频率学派特性且更具鲁棒性。
- 加权似然方法将NML框架扩展至未加权NML未定义的模型,利用其他比较或原假设的附加信息作为单一样本权重。
- 通过将偶然信息(如其他蛋白质或比较的数据)赋予一个观测值的权重,该方法确保了稳定性和鲁棒性。
- 通过简化似然模型应用该方法,保持了在不同样本量和比较数量下的可解释性与一致性。
- 通过渐近分析验证了理论性质,表明误导性证据的概率随样本量增加而趋于零。
实验结果
研究问题
- RQ1能否开发一种无需先验的证据度量,以支持简单原假设并避免p值的缺陷?
- RQ2如何使证据度量对其他比较中偶然信息的引入具有鲁棒性?
- RQ3所提出的方法是否能减少对传统多重比较校正的需求?
- RQ4区分信息能否可靠地指示对原假设的强证据,而不仅限于对备择假设的证据?
- RQ5在使用最少偶然信息的小样本量下,该方法表现如何?
主要发现
- 区分信息(DI)具有渐近可解释性,即随着样本量增加,误导性证据的概率收敛于零。
- 与p值不同,DI能够指示对简单原假设的强证据,解决了经典显著性检验的关键局限。
- 该方法具有极小化最大最优性,且无需对未观测结果进行平均,因此在频率学派设定下比贝叶斯因子更具鲁棒性。
- 将原假设作为单一样本权重时,DI值在视觉上与使用其他比较的偶然数据时几乎无法区分,即使在极小样本量(每组n=2)下也成立。
- 在每组n=4的20种蛋白质案例研究中,仅有一例蛋白质在原假设加权与偶然数据加权方法之间证据等级发生变化,表明其具有高度鲁棒性。
- 该方法在估计中表现出极小的后悔值,当使用最优信息时,后悔值保持恒定,而最大似然估计的后悔值则随样本均值变化极大。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。