Skip to main content
QUICK REVIEW

[论文解读] Selection Bias Correction and Eect Size Estimation under Dependence

Kean Ming Tan, Noah Simon|arXiv (Cornell University)|May 16, 2014
Statistical Methods and Inference参考文献 24被引用 3
一句话总结

本文提出了一种频率学派估计器,可在不假设检验统计量之间相互独立的情况下,校正大规模假设检验中效应量估计的选择偏差。该方法通过利用条件期望框架来调整选择引起的偏差,在模拟和真实基因表达数据中,尤其在存在依赖关系时,相较于以往方法显著提升了估计准确性。

ABSTRACT

We consider large-scale studies in which it is of interest to test a very large number of hypotheses, and then to estimate the eect sizes corresponding to the rejected hypotheses. For instance, this setting arises in the analysis of gene expression or DNA sequencing data. However, naive estimates of the eect sizes suer from selection bias, i.e., some of the largest naive estimates are large due to chance alone. Many authors have proposed methods to reduce the eects of selection bias under the assumption that the naive estimates of the eect sizes are independent. Unfortunately, when the eect size estimates are dependent, these existing techniques can have very poor performance, and in practice there will often be dependence. We propose an estimator that adjusts for selection bias under a recently-proposed frequentist framework, without the independence assumption. We study some properties of the proposed estimator, and illustrate that it outperforms past proposals in a simulation study and on two gene expression data sets.

研究动机与目标

  • 解决在基因组学和高通量数据分析中常见的大规模假设检验中效应量估计的选择偏差问题。
  • 克服现有偏差校正方法的局限性,这些方法通常假设效应量估计之间相互独立,而这一假设在真实数据中往往不成立。
  • 提出一种频率学派框架,用于效应量估计,能够考虑检验统计量之间的依赖关系,且无需强分布假设。
  • 提升大规模研究中被拒绝原假设的效应量估计准确性,尤其在选择偏差导致原始估计值被过度膨胀时。

提出的方法

  • 提出一种基于频率学派框架的偏差校正方法,通过条件化于选择事件,调整因较大效应量估计更可能被选中而产生的偏差。
  • 采用条件期望方法,通过考虑在依赖关系下检验统计量的联合分布,来估计真实效应量。
  • 使用多元正态近似来建模效应量估计之间的依赖结构,从而实现在估计值相关时的校正。
  • 通过在一般依赖结构下求解给定观测检验统计量和选择结果时的真实效应量的期望,推导出该估计器。
  • 避免依赖贝叶斯先验或重采样方法,使该方法适用于大规模生物数据分析中常见的频率学派设置。
  • 通过模拟研究和在两个基因表达数据集上的真实数据应用,验证估计器的性能。

实验结果

研究问题

  • RQ1当检验统计量存在依赖关系而非独立时,选择偏差如何影响效应量估计?
  • RQ2在不假设效应量估计之间相互独立的前提下,频率学派估计器能否在校正依赖关系下的选择偏差?
  • RQ3在现实依赖结构下,所提出的方法与现有偏差校正技术相比性能如何?
  • RQ4在具有相关检验统计量的真实世界基因表达数据中,所提出的估计器是否能保持良好的准确性和覆盖区间?

主要发现

  • 在依赖关系下,所提出的估计器显著降低了效应量估计的选择偏差,在模拟研究中优于基于独立性的方法。
  • 即使在检验数量庞大且检验统计量相关时,该方法仍能保持对真实效应量的准确覆盖。
  • 在两个真实基因表达数据集中,与原始估计和现有校正估计相比,所提出的估计器产生了更可靠且更不易被过度膨胀的效应量估计。
  • 该估计器在各种依赖结构下均表现稳健,包括高通量生物数据中常见的依赖模式。
  • 与以往方法相比,该方法在均方误差和偏差减少方面表现更优,尤其在检验统计量之间存在强相关性时更为显著。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。