Skip to main content
QUICK REVIEW

[论文解读] Asymptotic Results on Adaptive False Discovery Rate Controlling Procedures Based on Kernel Estimators

Pierre Neuvial|arXiv (Cornell University)|Mar 3, 2010
Statistical Methods in Clinical Trials参考文献 9被引用 13
一句话总结

本文研究了使用核密度估计器在 p 值为 1 处估计密度的自适应错误发现率(FDR)控制程序,从而提升渐近功效。证明了此类程序在保持正渐近功效的前提下,实现了更严格的 FDR 控制以及更广范围的目标 FDR 水平,尽管其收敛速率较慢,为非参数收敛速率 $m^{-k/(2k+1)}$,其中 $k$ 反映 p 值密度的光滑性。

ABSTRACT

The False Discovery Rate (FDR) is a commonly used type I error rate in multiple testing problems. It is defined as the expected False Discovery Proportion (FDP), that is, the expected fraction of false positives among rejected hypotheses. When the hypotheses are independent, the Benjamini-Hochberg procedure achieves FDR control at any pre-specified level. By construction, FDR control offers no guarantee in terms of power, or type II error. A number of alternative procedures have been developed, including plug-in procedures that aim at gaining power by incorporating an estimate of the proportion of true null hypotheses. In this paper, we study the asymptotic behavior of a class of plug-in procedures based on kernel estimators of the density of the $p$-values, as the number $m$ of tested hypotheses grows to infinity. In a setting where the hypotheses tested are independent, we prove that these procedures are asymptotically more powerful in two respects: (i) a tighter asymptotic FDR control for any target FDR level and (ii) a broader range of target levels yielding positive asymptotic power. We also show that this increased asymptotic power comes at the price of slower, non-parametric convergence rates for the FDP. These rates are of the form $m^{-k/(2k+1)}$, where $k$ is determined by the regularity of the density of the $p$-value distribution, or, equivalently, of the test statistics distribution. These results are applied to one- and two-sided tests statistics for Gaussian and Laplace location models, and for the Student model.

研究动机与目标

  • 研究使用核估计器在 1 处估计 p 值密度的自适应 FDR 控制程序的渐近行为。
  • 评估在检验数量不断增加的情况下,此类程序相较于标准 Benjamini-Hochberg 程序是否具有更高的功效。
  • 量化渐近功效提升与错误发现比例(FDP)收敛速率变慢之间的权衡。
  • 建立基于核的插补估计器对原假设比例 ($\hat{\pi}_0$) 实现有效 FDR 控制的条件,适用于大规模多重检验。
  • 推导这些程序下 FDP 的收敛速率,并将其与 p 值密度的光滑性联系起来。

提出的方法

  • 使用核估计器非参数地估计在 1 处的 p 值密度,该估计用于估计 $\pi_0$,即原假设的比例。
  • 在水平 $\alpha / \hat{\pi}_0$ 下应用 Benjamini-Hochberg 程序,以实现自适应 FDR 控制。
  • 在 $\hat{\pi}_0$ 以概率收敛于 $\pi_{0,\infty} \geq \pi_0$ 的假设下,使用 Delta 方法分析错误发现比例(FDP)的渐近分布。
  • 推导出 FDP 的收敛速率为 $m^{-k/(2k+1)}$,其中 $k$ 由 p 值密度的 Hölder 正则性决定,反映底层检验统计量分布的光滑性。
  • 在似然比 $f_1/f_0$ 在零附近满足正则性条件时,建立程序的一致性和纯洁性,尤其在对称模型(如正态分布和拉普拉斯分布)中成立。
  • 将结果应用于高斯、拉普拉斯和学生 t 分布中的单侧和双侧检验,验证 $g_1(t)$(备择假设下 p 值的密度)的正则性条件。

实验结果

研究问题

  • RQ1使用基于核的密度估计器估计 $\pi_0$ 是否能实现比标准 Benjamini-Hochberg 程序更紧的渐近 FDR 控制?
  • RQ2与标准 BH 程序相比,此类自适应程序是否能在更广范围的目标 FDR 水平下实现正的渐近功效?
  • RQ3在这些基于核的自适应程序下,错误发现比例(FDP)的收敛速率是多少?
  • RQ4p 值密度的光滑性(以 Hölder 正则性衡量)如何影响 FDP 的收敛速率?
  • RQ5在何种测试统计量分布条件下(如对称性、似然比的正则性),这些程序能维持 FDR 控制和渐近功效?

主要发现

  • 基于核的自适应 FDR 程序在任意目标 FDR 水平 $\alpha$ 下,均实现了比标准 Benjamini-Hochberg 程序更紧的渐近 FDR 控制。
  • 与标准 BH 程序相比,该程序在更广范围的目标 FDR 水平下表现出正的渐近功效。
  • 基于核的程序下,FDP 的收敛速率为 $m^{-k/(2k+1)}$,其中 $k$ 为 p 值密度的 Hölder 正则性指数,反映了更慢的非参数收敛速度。
  • FDP 的渐近分布为正态分布,其方差为 $w^2 = s_0^2 \pi_0^2 \alpha^2 / \pi_{0,\infty}^4$,该结果通过 Delta 方法推导得出。
  • 在对称模型(如正态分布、拉普拉斯分布)的双侧检验中,备择假设下 p 值的密度 $g_1(t)$ 在 $t=1$ 处可导,且在 $f_1/f_0$ 额外光滑性假设下二次可导。
  • 若似然比 $f_1/f_0$ 在零处可导,则 $g_1(t)$ 在 $t=1$ 处可导,这对 FDP 的渐近正态性至关重要。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。