Skip to main content
QUICK REVIEW

[论文解读] A leave-p-out based estimation of the proportion of null hypotheses

Alain Célisse, Stéphane Robin|ArXiv.org|Apr 8, 2008
Statistical Methods in Clinical Trials参考文献 27被引用 3
一句话总结

本文提出了一种基于直方图的留p-out交叉验证估计器,用于在多重检验中估计原假设真实比例($ olimits_0$),放松了强可识别性假设。该方法通过实现一种渐近控制FDR的插件式多重检验程序,提升了统计功效,在模拟中尤其在$ olimits_0$较小时优于Benjamini-Hochberg程序。

ABSTRACT

In the multiple testing context, a challenging problem is the estimation of the proportion $π_0$ of true-null hypotheses. A large number of estimators of this quantity rely on identifiability assumptions that either appear to be violated on real data, or may be at least relaxed. Under independence, we propose an estimator $\hatπ_0$ based on density estimation using both histograms and cross-validation. Due to the strong connection between the false discovery rate (FDR) and $π_0$, many multiple testing procedures (MTP) designed to control the FDR may be improved by introducing an estimator of $π_0$. We provide an example of such an improvement (plug-in MTP) based on the procedure of Benjamini and Hochberg. Asymptotic optimality results may be derived for both $\hatπ_0$ and the resulting plug-in procedure. The latter ensures the desired asymptotic control of the FDR, while it is more powerful than the BH-procedure. Finally, we compare our estimator of $π_0$ with other widespread estimators in a wide range of simulations. We obtain better results than other tested methods in terms of mean square error (MSE) of the proposed estimator. Finally, both asymptotic optimality results and the interest in tightly estimating $π_0$ are confirmed (empirically) by results obtained with the plug-in MTP.

研究动机与目标

  • 解决在现有估计器依赖强且常被违反的可识别性假设时,估计多重检验场景中原假设真实比例$ olimits_0$的挑战。
  • 开发一种灵活、完全自适应的$ olimits_0$估计器,即使在真实数据中p值在1附近被人为放大时仍保持可靠性。
  • 通过插入准确的$ olimits_0$估计值,提升FDR控制程序的统计功效,同时确保渐近FDR控制。
  • 提供一种计算高效的非参数方法,使用非规则直方图与留p-out交叉验证进行$ olimits_0$估计,避免用户指定的调优参数。

提出的方法

  • 该方法使用非规则直方图估计p值密度,从而在建模复杂p值分布(包括U形密度)时具有灵活性。
  • 采用留p-out交叉验证(LPO)选择最优直方图分箱,确保数据驱动的自适应带宽选择,无需用户定义参数。
  • 在放宽的可识别性假设下,基于靠近1处的直方图密度估计,推导出$ olimits_0$的估计量,即原假设下p值所占比例的估计。
  • 通过将Benjamini-Hochberg程序中的未知$ olimits_0$替换为$ olimits_0$,构建插件式多重检验程序(plug-in MTP),在保持FDR控制的同时提升统计功效。
  • 在独立性和温和正则性条件下,证明了$ olimits_0$的渐近一致性以及插件MTP的渐近FDR控制。
  • 该方法避免依赖参数模型或固定调优参数,因此对1附近均匀性假设的违反具有鲁棒性。

实验结果

研究问题

  • RQ1能否开发一种非参数、数据驱动的$ olimits_0$估计器,使其在p值在1附近被人为放大、违反标准可识别性假设时仍保持有效?
  • RQ2与固定参数方法相比,留p-out交叉验证是否在$ olimits_0$估计的直方图密度估计中提供了更优、更自适应的方法?
  • RQ3使用估计的$ olimits_0$的插件式多重检验程序能否在仍渐近控制FDR的同时,实现比Benjamini-Hochberg程序更高的统计功效?
  • RQ4在多样化的模拟情景下,所提出的$ olimits_0$估计器与现有方法相比,其均方误差(MSE)表现如何?
  • RQ5在留p-out交叉验证中选择$p > 1$的影响是什么?与标准的留一法相比,其是否能提升估计精度?

主要发现

  • 在广泛多样的模拟情景中,所提出的$ olimits_0$估计器的均方误差(MSE)低于其他广泛使用的估计器,表明其具有更优的偏差-方差权衡。
  • 实证结果表明,基于$ olimits_0$的插件多重检验程序在有限样本中能保持渐近FDR控制。
  • 当$ olimits_0$较小时,插件MTP始终比Benjamini-Hochberg程序更具统计功效,原因在于对原假设比例的估计更准确。
  • 当$p > 1$时,留p-out交叉验证(LPO)的性能优于留一法(LOO),且其性能几乎与已知$ olimits_0$的“理想”程序相当。
  • 在p值分布呈“U形”的情况下,该估计器依然具有鲁棒性,而许多现有方法在此类情形下会失效,原因在于其放宽了可识别性假设。
  • 该方法完全自适应,无需用户指定参数,而与之相比,Schweder和Spjøtvoll估计器依赖于固定的$ olimits$。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。