Skip to main content
QUICK REVIEW

[Paper Review] FarmTest: Factor-Adjusted Robust Multiple Testing With Approximate False Discovery Control

Jianqing Fan, Yuan Ke|arXiv (Cornell University)|Nov 15, 2017
Statistical Methods in Clinical Trials56 references4 citations
TL;DR

This paper proposes FarmTest, a factor-adjusted robust multiple testing procedure that controls the false discovery proportion (FDP) under general dependence and heavy-tailed data. By integrating robust estimation via Huber loss and latent factor modeling, FarmTest achieves consistent FDP estimation and improved power, especially under non-normal, correlated high-dimensional data.

ABSTRACT

Large-scale multiple testing with correlated and heavy-tailed data arises in a wide range of research areas from genomics, medical imaging to finance. Conventional methods for estimating the false discovery proportion (FDP) often ignore the effect of heavy-tailedness and the dependence structure among test statistics, and thus may lead to inefficient or even inconsistent estimation. Also, the commonly imposed joint normality assumption is arguably too stringent for many applications. To address these challenges, in this article we propose a factor-adjusted robust multiple testing (FarmTest) procedure for large-scale simultaneous inference with control of the FDP. We demonstrate that robust factor adjustments are extremely important in both controlling the FDP and improving the power. We identify general conditions under which the proposed method produces consistent estimate of the FDP. As a byproduct that is of independent interest, we establish an exponential-type deviation inequality for a robust <i>U</i>-type covariance estimator under the spectral norm. Extensive numerical experiments demonstrate the advantage of the proposed method over several state-of-the-art methods especially when the data are generated from heavy-tailed distributions. The proposed procedures are implemented in the R-package FarmTest. Supplementary materials for this article are available online.

Motivation & Objective

  • Address the limitations of conventional multiple testing methods that assume independence and joint normality, which can lead to inconsistent FDP estimation under dependence and heavy-tailed distributions.
  • Develop a method that accounts for latent factor structures in high-dimensional test statistics, common in genomics, neuroscience, and finance.
  • Achieve consistent estimation of the false discovery proportion (FDP) under weak or strong dependence and non-normal error distributions.
  • Improve statistical power by adjusting for common factors while maintaining robustness to heavy-tailed errors.
  • Establish theoretical guarantees for FDP control under general conditions, including a novel exponential-type deviation inequality for robust U-statistic covariance estimators.

Proposed method

  • Model high-dimensional test statistics using a factor model with latent factors and idiosyncratic errors, where the factor structure captures general dependence.
  • Apply Huber loss-based estimation to robustly estimate factor loadings and common factors, reducing sensitivity to heavy-tailed errors.
  • Use a robust U-type covariance estimator for the idiosyncratic component, with a new exponential-type deviation bound under the spectral norm.
  • Adjust test statistics by removing estimated common factors to reduce dependence, enabling more accurate FDP control.
  • Control the false discovery proportion (FDP) via a step-up procedure based on factor-adjusted p-values, with theoretical consistency guarantees.
  • Implement the method in an R package called FarmTest for practical application in large-scale inference.

Experimental results

Research questions

  • RQ1Can robust factor adjustment improve FDP estimation and control under general dependence and heavy-tailed error distributions?
  • RQ2Under what conditions is the proposed FDP estimator consistent when the data are non-normal and dependent?
  • RQ3How does the robustness of the Huber-based estimation compare to classical methods under heavy-tailed and correlated data?
  • RQ4What is the impact of latent factor structures on the power and accuracy of multiple testing procedures in high-dimensional settings?
  • RQ5Can a robust covariance estimator with strong concentration properties be developed for use in high-dimensional factor models?

Key findings

  • FarmTest achieves consistent estimation of the false discovery proportion (FDP) under general dependence and heavy-tailed distributions, even when the joint normality assumption fails.
  • The method significantly improves statistical power compared to state-of-the-art competitors, especially under heavy-tailed error distributions such as t-distribution with low degrees of freedom.
  • The robust U-type covariance estimator achieves an exponential-type deviation bound under the spectral norm, enabling strong theoretical guarantees.
  • Empirical results show that FarmTest maintains accurate FDP control across various sample sizes and error distributions, outperforming existing methods in both FDR control and power.
  • Theoretical analysis confirms that factor adjustment is essential for consistent FDP estimation, and robustness via Huber loss is critical when errors deviate from normality.
  • The R package FarmTest enables practical implementation of the method, supporting reproducible large-scale inference in genomics, neuroscience, and finance.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.