Skip to main content
QUICK REVIEW

[Paper Review] Control of the False Discovery Rate Under Arbitrary Covariance Dependence

Xu Han, Weijie Gu|arXiv (Cornell University)|Dec 20, 2010
Statistical Methods in Clinical Trials25 references15 citations
TL;DR

This paper proposes a Principal Factor Approximation (PFA) method to control the False Discovery Rate (FDR) under arbitrary dependence structures in high-dimensional multiple testing, by removing strong common factors via spectral decomposition of the covariance matrix. The method enables consistent estimation of the False Discovery Proportion (FDP) and outperforms existing approaches like Efron's under correlated test statistics.

ABSTRACT

Multiple hypothesis testing is a fundamental problem in high dimensional inference, with wide applications in many scientific fields. In genome-wide association studies, tens of thousands of tests are performed simultaneously to find if any genes are associated with some traits and those tests are correlated. When test statistics are correlated, false discovery control becomes very challenging under arbitrary dependence. In the current paper, we propose a new methodology based on principal factor approximation, which successfully substracts the common dependence and weakens significantly the correlation structure, to deal with an arbitrary dependence structure. We derive the theoretical distribution for false discovery proportion (FDP) in large scale multiple testing when a common threshold is used and provide a consistent FDP. This result has important applications in controlling FDR and FDP. Our estimate of FDP compares favorably with Efron (2007)'s approach, as demonstrated by in the simulated examples. Our approach is further illustrated by some real data applications.

Motivation & Objective

  • To address the challenge of controlling FDR when test statistics are arbitrarily dependent, a common issue in high-dimensional inference.
  • To develop a method that fully incorporates covariance information rather than assuming independence or specific dependence structures.
  • To derive a consistent estimator of the False Discovery Proportion (FDP) under general dependence, enabling reliable FDR control.
  • To improve upon existing FDR procedures—such as Benjamini-Hochberg and Storey’s method—that lose accuracy under high correlation.
  • To provide a theoretically grounded, scalable solution applicable to real-world problems like genome-wide association studies (GWAS).

Proposed method

  • Uses spectral decomposition of the known covariance matrix $\Sigma$ to extract the principal factors responsible for strong dependence among test statistics.
  • Applies principal factor approximation (PFA) to remove the dominant common factors, thereby weakening the overall dependence structure.
  • Derives the asymptotic distribution of the False Discovery Proportion (FDP) under large $p$, accounting for residual weak dependence after factor removal.
  • Employs a threshold-based procedure where $\widehat{\text{FDR}}(t) = \widehat{p}_0 t / (R(t) \vee 1)$, solving for $t$ such that $\widehat{\text{FDR}}(t) \leq \alpha$, with $\widehat{p}_0$ estimated from the data.
  • Uses the normal cumulative distribution function $\Phi$ and standard normal density $\phi$ to model tail probabilities in the FDP estimation, incorporating estimated factor loadings and weights.
  • Applies the Cauchy-Schwarz inequality and concentration bounds to establish asymptotic consistency of the FDP estimator under regularity conditions on eigenvalues and factor loadings.

Experimental results

Research questions

  • RQ1Can FDR be consistently controlled when test statistics exhibit arbitrary dependence structures, including high correlation?
  • RQ2How can the common dependence structure in high-dimensional data be effectively removed to improve FDR estimation?
  • RQ3What is the theoretical distribution of the False Discovery Proportion (FDP) under general dependence, and can it be consistently estimated?
  • RQ4How does the proposed PFA method compare in performance to Efron’s (2007) empirical Bayes approach under correlated test statistics?
  • RQ5To what extent does the method maintain power and accuracy in real-world applications such as genome-wide association studies?

Key findings

  • The proposed PFA method achieves consistent estimation of the False Discovery Proportion (FDP) under arbitrary dependence, even when the covariance matrix is non-trivial.
  • Theoretical analysis shows that the FDP estimator converges in probability to the true FDP under regularity conditions, including bounded eigenvalues and factor loadings.
  • The method outperforms Efron’s (2007) approach in simulated examples, particularly in high-correlation settings where FDR control is most compromised.
  • The FDP estimator is shown to be $O_p(\sqrt{k/m})$-consistent, where $k$ is the number of factors and $m$ is the sample size, under appropriate moment conditions.
  • The procedure maintains control of FDR at the nominal level $\alpha$ even when test statistics are highly correlated, as demonstrated in both simulations and real data applications.
  • The method is scalable and applicable to large-scale problems such as GWAS, where tens of thousands of hypotheses are tested simultaneously with complex dependence structures.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.