Skip to main content
QUICK REVIEW

[Paper Review] A leave-p-out based estimation of the proportion of null hypotheses

Alain Célisse, Stéphane Robin|ArXiv.org|Apr 8, 2008
Statistical Methods in Clinical Trials27 references3 citations
TL;DR

This paper proposes a leave-p-out cross-validated histogram-based estimator for the proportion of true null hypotheses ($\pi_0$) in multiple testing, relaxing strong identifiability assumptions. The method improves power in FDR-controlling procedures by enabling a plug-in multiple testing procedure that asymptotically controls the FDR while outperforming the Benjamini-Hochberg procedure in simulations, especially when $\pi_0$ is small.

ABSTRACT

In the multiple testing context, a challenging problem is the estimation of the proportion $π_0$ of true-null hypotheses. A large number of estimators of this quantity rely on identifiability assumptions that either appear to be violated on real data, or may be at least relaxed. Under independence, we propose an estimator $\hatπ_0$ based on density estimation using both histograms and cross-validation. Due to the strong connection between the false discovery rate (FDR) and $π_0$, many multiple testing procedures (MTP) designed to control the FDR may be improved by introducing an estimator of $π_0$. We provide an example of such an improvement (plug-in MTP) based on the procedure of Benjamini and Hochberg. Asymptotic optimality results may be derived for both $\hatπ_0$ and the resulting plug-in procedure. The latter ensures the desired asymptotic control of the FDR, while it is more powerful than the BH-procedure. Finally, we compare our estimator of $π_0$ with other widespread estimators in a wide range of simulations. We obtain better results than other tested methods in terms of mean square error (MSE) of the proposed estimator. Finally, both asymptotic optimality results and the interest in tightly estimating $π_0$ are confirmed (empirically) by results obtained with the plug-in MTP.

Motivation & Objective

  • To address the challenge of estimating $\pi_0$, the proportion of true null hypotheses, in multiple testing scenarios where existing estimators rely on strong, often violated identifiability assumptions.
  • To develop a flexible, fully adaptive estimator of $\pi_0$ that remains reliable even when p-values are artificially inflated near 1, as observed in real data.
  • To improve the power of false discovery rate (FDR)-controlling procedures by plugging in an accurate $\pi_0$ estimate, while ensuring asymptotic FDR control.
  • To provide a computationally efficient, nonparametric approach to $\pi_0$ estimation using non-regular histograms and leave-p-out cross-validation, avoiding user-specified tuning parameters.

Proposed method

  • The method estimates the density of p-values using non-regular histograms, which allows flexibility in modeling complex p-value distributions, including U-shaped densities.
  • Leave-p-out cross-validation (LPO) is used to select the optimal histogram binning, ensuring data-driven, adaptive bandwidth selection without requiring user-defined parameters.
  • The estimator $\widehat{\pi}_0$ is derived as the estimated proportion of p-values under the null, based on the histogram density estimate near 1, under relaxed identifiability assumptions.
  • A plug-in multiple testing procedure (plug-in MTP) is constructed by replacing the unknown $\pi_0$ in the Benjamini-Hochberg procedure with $\widehat{\pi}_0$, improving power while maintaining FDR control.
  • Asymptotic consistency of $\widehat{\pi}_0$ and asymptotic FDR control of the plug-in MTP are established under independence and mild regularity conditions.
  • The method avoids reliance on parametric models or fixed tuning parameters, making it robust to violations of the uniformity assumption near 1.

Experimental results

Research questions

  • RQ1Can a nonparametric, data-driven estimator of $\pi_0$ be developed that remains valid when p-values are artificially inflated near 1, violating standard identifiability assumptions?
  • RQ2Does leave-p-out cross-validation provide a superior, adaptive method for histogram-based density estimation in the context of $\pi_0$ estimation compared to fixed-parameter methods?
  • RQ3Can a plug-in multiple testing procedure using the estimated $\pi_0$ achieve better power than the Benjamini-Hochberg procedure while still asymptotically controlling the FDR?
  • RQ4How does the performance of the proposed $\pi_0$ estimator compare to existing methods in terms of mean squared error (MSE) across diverse simulation scenarios?
  • RQ5What is the impact of choosing $p > 1$ in leave-p-out cross-validation, and does it improve estimation accuracy compared to the standard leave-one-out approach?

Key findings

  • The proposed $\pi_0$ estimator achieves lower mean squared error (MSE) than other widely used estimators across a wide range of simulation scenarios, indicating superior bias-variance tradeoff.
  • The plug-in multiple testing procedure based on $\widehat{\pi}_0$ is empirically shown to maintain asymptotic control of the false discovery rate (FDR) in finite samples.
  • The plug-in MTP is consistently more powerful than the Benjamini-Hochberg procedure, especially when $\pi_0$ is small, due to more accurate estimation of the null proportion.
  • Leave-p-out cross-validation (LPO) with $p > 1$ yields better performance than leave-one-out (LOO), and LPO is nearly as powerful as the oracle procedure that knows $\pi_0$ exactly.
  • The estimator remains robust in the 'U-shape' p-value distribution case, where many existing methods fail, due to its relaxed identifiability assumption.
  • The method is fully adaptive and does not require user-specified parameters, unlike the Schweder and Spjøtvoll estimator, which depends on a fixed $\lambda$.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.