Skip to main content
QUICK REVIEW

[Paper Review] Selection Bias Correction and Effect Size Estimation under Dependence

Kean Ming Tan, Noah Simon|arXiv (Cornell University)|May 16, 2014
Statistical Methods and Inference4 citations
TL;DR

This paper proposes a novel frequentist approach to correct selection bias in high-dimensional effect size estimation when unadjusted test statistics are dependent, using nonparametric and parametric bootstrap resampling to account for correlation structures. The method outperforms existing independent-estimates-based techniques in simulations and real gene expression data, especially under dependence, without requiring normality assumptions.

ABSTRACT

We consider large-scale studies in which it is of interest to test a very large number of hypotheses, and then to estimate the effect sizes corresponding to the rejected hypotheses. For instance, this setting arises in the analysis of gene expression or DNA sequencing data. However, naive estimates of the effect sizes suffer from selection bias, i.e., some of the largest naive estimates are large due to chance alone. Many authors have proposed methods to reduce the effects of selection bias under the assumption that the naive estimates of the effect sizes are independent. Unfortunately, when the effect size estimates are dependent, these existing techniques can have very poor performance, and in practice there will often be dependence. We propose an estimator that adjusts for selection bias under a recently-proposed frequentist framework, without the independence assumption. We study some properties of the proposed estimator, and illustrate that it outperforms past proposals in a simulation study and on two gene expression data sets.

Motivation & Objective

  • Address selection bias in high-dimensional effect size estimation when unadjusted estimates are dependent, a common issue in genomics and large-scale inference.
  • Overcome the limitations of existing methods that assume independence among unadjusted estimates, which is often violated in real-world data like gene expression.
  • Develop a robust, nonparametric framework that avoids strong parametric assumptions (e.g., normality) on the data or test statistics.
  • Provide a practical and generalizable method applicable to a wide range of test statistics and data types, including non-normal distributions.
  • Demonstrate improved accuracy in effect size estimation through simulation and real data applications on prostate and lung cancer gene expression datasets.

Proposed method

  • Adapt the frequentist framework of Simon and Simon (2016) to correct selection bias under dependence, without assuming independence among unadjusted estimates.
  • Use a nonparametric bootstrap to resample the joint distribution of test statistics, preserving observed dependence structures in the resampled data.
  • Use a parametric bootstrap under a multivariate normal assumption to model the dependence structure of the unadjusted estimates.
  • Estimate the conditional expectation of the true effect size given the observed test statistic by averaging over bootstrap resamples, adjusting for selection bias.
  • Apply the bootstrap-based estimator to correct the largest effect size estimates, which are most prone to overestimation due to selection bias.
  • Validate the method via cross-validation with training/test splits, measuring performance via sum of squared differences between corrected estimates and unadjusted estimates on the test set.

Experimental results

Research questions

  • RQ1How does dependence among unadjusted test statistics affect the accuracy of existing selection bias correction methods that assume independence?
  • RQ2Can a bootstrap-based frequentist framework effectively correct selection bias when unadjusted estimates are dependent?
  • RQ3How does the performance of the proposed method compare to empirical Bayes and James-Stein-type estimators under dependence?
  • RQ4Does accounting for dependence via bootstrap resampling lead to more accurate effect size estimates in real gene expression data?
  • RQ5To what extent do the proposed methods reduce overestimation of extreme effect sizes in high-dimensional settings with correlated features?

Key findings

  • The proposed parametric and nonparametric bootstrap methods significantly outperform existing approaches like tweedie, nlpden, and James-Stein when unadjusted estimates are dependent.
  • In the prostate cancer data set, the parametric bootstrap with correlation correction (para-cor) achieved the lowest sum of squared differences (51.07 for k=15), outperforming all other methods.
  • In the lung cancer data set, para-cor again showed the best performance (59.38 for k=15), indicating robustness to dependence in real-world data.
  • The nonparametric bootstrap (nonpara) performed nearly as well as the parametric version, suggesting it is effective even without normality assumptions.
  • Methods relying on density estimation (nlpden, tweedie) performed poorly due to insufficient extreme values for accurate marginal density estimation.
  • The parametric bootstrap with correlation correction (para-cor) outperformed the uncorrelated version (para-uncor), confirming that modeling dependence improves accuracy.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.