[论文解读] FARM-Test: Factor-Adjusted Robust Multiple Testing with False Discovery Control
该论文提出FARM-Test,一种因子调整的稳健多重假设检验方法,在一般依赖结构和重尾分布下控制错误发现比例(FDP)。通过整合基于U-统计量的稳健协方差估计与因子模型,该方法实现了FDP估计的一致性,并在非正态、依赖数据设置下显著提升了检验功效。
Large-scale multiple testing with correlated and heavy-tailed data arises in a wide range of research areas from genomics, medical imaging to finance. Conventional methods for estimating the false discovery proportion (FDP) often ignore the effect of heavy-tailedness and the dependence structure among test statistics, and thus may lead to inefficient or even inconsistent estimation. Also, the assumption of joint normality is often imposed, which is too stringent for many applications. To address these challenges, in this paper we propose a factor-adjusted robust procedure for large-scale simultaneous inference with control of the false discovery proportion. We demonstrate that robust factor adjustments are extremely important in both improving the power of the tests and controlling FDP. We identify general conditions under which the proposed method produces consistent estimate of the FDP. As a byproduct that is of independent interest, we establish an exponential-type deviation inequality for a robust $U$-type covariance estimator under the spectral norm. Extensive numerical experiments demonstrate the advantage of the proposed method over several state-of-the-art methods especially when the data are generated from heavy-tailed distributions. Our proposed procedures are implemented in the {\sf R}-package {\sf FarmTest}.
研究动机与目标
- 解决传统多重假设检验方法假设联合正态性、忽略依赖性和重尾性的局限性。
- 开发一种在检验统计量存在依赖性和重尾性时仍能保持准确错误发现比例(FDP)控制的方法。
- 通过调整潜在因子并稳健化协方差估计,提升大规模推断中的统计功效。
- 在超越正态性的广义条件下,建立FDP估计的一致性理论。
- 提供一个实用、可实现的R软件包(FarmTest),用于基因组学、金融和医学影像等实际应用。
提出的方法
- 使用稳健的U型协方差估计器处理重尾分布,并为其谱范数建立了新的指数型偏差不等式。
- 应用因子模型通过去除公共因子来解释检验统计量之间的普遍依赖结构。
- 通过在稳健协方差矩阵上进行主成分分析,估计并去除潜在因子的影响,从而对检验统计量进行调整。
- 采用在弱依赖和重尾性下一致的错误发现比例(FDP)估计器。
- 结合因子调整与稳健推断,在保持FDP控制的同时提升检验功效。
- 推导出在数据偏离正态性时,FDP估计器仍保持一致性的理论条件。
实验结果
研究问题
- RQ1在数据存在依赖性和重尾性时,如何在高维多重假设检验中一致地估计错误发现比例(FDP)?
- RQ2在非正态和依赖数据条件下,因子调整在多大程度上能提升多重假设检验方法的统计功效和准确性?
- RQ3在不假设正态性的情况下,稳健协方差估计器能否在重尾分布下提供可靠的推断?
- RQ4因子调整的稳健方法在何种理论条件下能实现一致的FDP估计?
- RQ5在存在重尾和依赖性的设置下,所提出方法与现有最先进方法相比在实证表现上如何?
主要发现
- 所提出的FARM-Test方法在一般依赖结构和重尾分布下,即使正态性假设不成立,也能实现错误发现比例(FDP)的一致估计。
- 与传统方法相比,稳健因子调整显著提升了统计功效,尤其在重尾设置下表现更优。
- 针对稳健U型协方差估计器的谱范数,建立了指数型偏差不等式,从而提供了有限样本的理论保证。
- 在广泛的数值实验中,该方法优于现有最先进方法,尤其在数据呈现重尾性和依赖性时表现更佳。
- FARM-Test方法已通过R软件包{\tt FarmTest}实现,支持在基因组学、医学影像和金融数据分析中的实际应用。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。