[论文解读] FarmTest: Factor-Adjusted Robust Multiple Testing With Approximate False Discovery Control
本文提出 FarmTest,一种经过因子调整的稳健多重假设检验方法,在一般依赖结构和重尾数据下控制错误发现比例(FDP)。通过整合基于 Huber 损失的稳健估计与潜在因子模型,FarmTest 实现了 FDP 的一致估计,并在非正态、相关性高的高维数据中显著提升了检验效能。
Large-scale multiple testing with correlated and heavy-tailed data arises in a wide range of research areas from genomics, medical imaging to finance. Conventional methods for estimating the false discovery proportion (FDP) often ignore the effect of heavy-tailedness and the dependence structure among test statistics, and thus may lead to inefficient or even inconsistent estimation. Also, the commonly imposed joint normality assumption is arguably too stringent for many applications. To address these challenges, in this article we propose a factor-adjusted robust multiple testing (FarmTest) procedure for large-scale simultaneous inference with control of the FDP. We demonstrate that robust factor adjustments are extremely important in both controlling the FDP and improving the power. We identify general conditions under which the proposed method produces consistent estimate of the FDP. As a byproduct that is of independent interest, we establish an exponential-type deviation inequality for a robust <i>U</i>-type covariance estimator under the spectral norm. Extensive numerical experiments demonstrate the advantage of the proposed method over several state-of-the-art methods especially when the data are generated from heavy-tailed distributions. The proposed procedures are implemented in the R-package FarmTest. Supplementary materials for this article are available online.
研究动机与目标
- 解决传统多重检验方法假设独立性和联合正态性所带来的局限,这些假设在依赖性和重尾分布下会导致 FDP 估计不一致。
- 开发一种能够考虑高维检验统计量中潜在因子结构的方法,此类结构在基因组学、神经科学和金融学中普遍存在。
- 在弱依赖或强依赖以及非正态误差分布下,实现错误发现比例(FDP)的一致估计。
- 通过调整共同因子来提升统计效能,同时保持对重尾误差的稳健性。
- 在一般条件下建立 FDP 控制的理论保证,包括针对稳健 U-统计量协方差估计器的新型指数型尾部不等式。
提出的方法
- 使用包含潜在因子和特异性误差的因子模型来建模高维检验统计量,其中因子结构捕捉整体依赖性。
- 应用基于 Huber 损失的估计方法,稳健估计因子载荷和共同因子,降低对重尾误差的敏感性。
- 对特异性成分使用稳健的 U-型协方差估计器,并在谱范数下建立新的指数型偏差界。
- 通过去除估计的共同因子来调整检验统计量,以减少依赖性,从而实现更精确的 FDP 控制。
- 基于因子调整后的 p 值,采用逐步提升程序控制错误发现比例(FDP),并具备理论一致性保证。
- 在 R 语言中实现该方法,发布为名为 FarmTest 的 R 包,以支持在大规模推断中的实际应用。
实验结果
研究问题
- RQ1在一般依赖和重尾误差分布下,稳健的因子调整能否改善 FDP 估计与控制?
- RQ2当数据为非正态且存在依赖时,所提出的 FDP 估计器在何种条件下具有一致性?
- RQ3在重尾且相关的数据中,基于 Huber 的估计方法与经典方法相比,其稳健性如何?
- RQ4潜在因子结构对高维设定下多重检验方法的效能和准确性有何影响?
- RQ5能否为高维因子模型开发一种具有强集中性质的稳健协方差估计器?
主要发现
- FarmTest 在一般依赖和重尾分布下实现了错误发现比例(FDP)的一致估计,即使在联合正态性假设不成立时亦成立。
- 与现有最先进方法相比,该方法在重尾误差分布(如自由度较低的 t 分布)下显著提升了统计效能。
- 稳健的 U-型协方差估计器在谱范数下实现了指数型偏差界,从而提供了强有力的理论保证。
- 实证结果表明,FarmTest 在各种样本大小和误差分布下均能保持准确的 FDP 控制,且在 FDR 控制和效能方面均优于现有方法。
- 理论分析证实,因子调整对于一致的 FDP 估计至关重要,而通过 Huber 损失实现的稳健性在误差偏离正态性时尤为关键。
- R 包 FarmTest 支持该方法的实际应用,可实现基因组学、神经科学和金融学中可复现的大规模推断。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。