[论文解读] Phase Transition and Regularized Bootstrap in Large Scale $t$-tests with False Discovery Rate Control
本文研究在使用正态分布或 t 分布近似 p 值时,大规模 t 检验中错误发现率(FDR)控制的有效性。当 log m ≈ c₀n¹ᐟ³ 时,这些近似方法失效,出现相变现象。本文提出一种正则化自展法,在仅具有有限六阶矩的重尾分布下仍能保持 FDR 控制,且在模拟中优于标准自展法。
Applying Benjamini and Hochberg (B-H) method to multiple Student's $t$ tests is a popular technique in gene selection in microarray data analysis. Because of the non-normality of the population, the true p-values of the hypothesis tests are typically unknown. Hence, it is common to use the standard normal distribution N(0,1), Student's $t$ distribution $t_{n-1}$ or the bootstrap method to estimate the p-values. In this paper, we first study N(0,1) and $t_{n-1}$ calibrations. We prove that, when the population has the finite 4-th moment and the dimension $m$ and the sample size $n$ satisfy $\log m=o(n^{1/3})$, B-H method controls the false discovery rate (FDR) at a given level $α$ asymptotically with p-values estimated from N(0,1) or $t_{n-1}$ distribution. However, a phase transition phenomenon occurs when $\log m\geq c_{0}n^{1/3}$. In this case, the FDR of B-H method may be larger than $α$ or even tends to one. In contrast, the bootstrap calibration is accurate for $\log m=o(n^{1/2})$ as long as the underlying distribution has the sub-Gaussian tails. However, such light tailed condition can not be weakened in general. The simulation study shows that for the heavy tailed distributions, the bootstrap calibration is very conservative. In order to solve this problem, a regularized bootstrap correction is proposed and is shown to be robust to the tails of the distributions. The simulation study shows that the regularized bootstrap method performs better than the usual bootstrap method.
研究动机与目标
- 分析在大规模多重检验中,使用标准正态分布或 t 分布近似 p 值时 FDR 控制的渐近行为。
- 识别在高维设置下,由于高维相变导致这些近似方法失效的条件。
- 开发一种稳健的自展校准替代方法,以在重尾分布下保持 FDR 控制。
- 提出并从理论上证明一种正则化自展方法,仅需有限六阶矩,优于标准自展法的保守性。
提出的方法
- 对大规模 t 检验中 p 值使用正态分布和 t 分布近似时的 FDR 控制进行理论分析。
- 基于维度 m 与样本量 n 的关系,推导相变阈值,特别是当 log m ≥ c₀n¹ᐟ³ 时。
- 使用自展校准以提高 p 值估计的准确性,尤其在次高斯尾部下表现更优。
- 引入一种正则化自展方法,通过截断极端观测值以降低对重尾的敏感性。
- 应用浓度不等式和矩界,证明截断后经验矩的收敛性。
- 从理论上证明正则化自展法下的 FDR 控制,表明当 log m = o(n¹ᐟ²) 时具有渐近有效性。
实验结果
研究问题
- RQ1在何种条件下,使用标准正态分布和 t 分布近似 p 值时,无法在大规模 t 检验中控制 FDR?
- RQ2当 log m ≥ c₀n¹ᐟ³ 时,FDR 控制的相变性质及其阈值是什么?
- RQ3自展校准的性能如何依赖于尾部行为?能否在重尾分布下进一步改进?
- RQ4正则化自展方法是否能在仅具有有限六阶矩的重尾分布下保持 FDR 控制?
- RQ5在具有重尾数据的有限样本模拟中,正则化自展法是否比标准自展法更具稳健性?
主要发现
- 当 log m = o(n¹ᐟ³) 时,在有限四阶矩下,使用标准正态分布或 t_{n-1} 分布近似 p 值,Benjamini-Hochberg 方法的 FDR 控制具有渐近有效性。
- 当 log m ≥ c₀n¹ᐟ³ 时发生相变;在非零偏度下,FDR 可能超过 α,甚至随着 log m/n¹ᐟ³ → ∞ 而趋于 1。
- 当底层分布具有次高斯尾部时,自展校准可保持 FDR 控制,条件是 log m = o(n¹ᐟ²)。
- 即使矩有限,标准自展法在重尾分布下仍过于保守。
- 所提出的正则化自展方法在仅具有有限六阶矩下,于 log m = o(n¹ᐟ²) 时确保 FDR 控制,且在模拟中优于标准自展法。
- 理论与模拟结果共同表明,正则化自展法对尾部厚重性具有稳健性,并在重尾设定下提供比标准自展法更精确的 FDR 控制。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。