Skip to main content
QUICK REVIEW

[论文解读] Accounting for unobserved covariates with varying degrees of estimability in high dimensional biological data

Chris McKennan, Dan L. Nicolae|arXiv (Cornell University)|Jan 3, 2018
Statistical Methods and Inference被引用 3
一句话总结

本文提出了一种针对高维生物数据的偏差校正估计器,可处理具有不同可估性的未观测协变量,纠正现有方法中的渐近偏差。该文证明,当所有协变量均可观测时,新估计器的渐近分布与普通最小二乘法(OLS)相同;并通过DNA甲基化数据表明,该方法更准确地估计了哮喘的直接效应,揭示了细胞组成在甲基化变异中起主要中介作用。

ABSTRACT

An important phenomenon in high dimensional biological data is the presence of unobserved covariates that can have a significant impact on the measured response. When these factors are also correlated with the covariate(s) of interest (i.e. disease status), ignoring them can lead to increased type I error and spurious false discovery rate estimates. We show that depending on the strength of this correlation and the informativeness of the observed data for the latent factors, previously proposed estimators for the effect of the covariate of interest that attempt to account for unobserved covariates are asymptotically biased, which corroborates previous practitioners' observations that these estimators tend to produce inflated test statistics. We then provide an estimator that corrects the bias and prove it has the same asymptotic distribution as the ordinary least squares estimator when every covariate is observed. Lastly, we use previously published DNA methylation data to show our method can more accurately estimate the direct effect of asthma on methylation than previously published methods, which underestimate the correlation between asthma and latent cell type heterogeneity. Our re-analysis shows that the majority of the variability in methylation due to asthma in those data is actually mediated through cell composition.

研究动机与目标

  • 解决未观测协变量(如细胞类型异质性或批次效应)对高维生物数据中效应估计产生偏差的问题。
  • 识别并校正当未观测协变量仅部分可从数据中估计时,现有估计器产生的渐近偏差。
  • 开发一种方法,即使未观测因素估计不佳,也能保持与完全观测到这些因素时相当的统计功效。
  • 为具有未观测混杂因子的高维因子模型提供一个理论基础坚实的推断框架。

提出的方法

  • 提出一种新估计器,通过两阶段程序校正普通最小二乘法(OLS)中的偏差,以纠正未观测协变量的影响。
  • 对残差进行因子分析以估计潜在因子,然后将朴素的OLS估计值对这些因子进行回归,以估计偏差分量。
  • 推导校正后估计器的渐近分布,并证明在完全观测到混杂因子时,其分布与OLS相同。
  • 运用高维随机矩阵理论分析偏差校正项的方差,表明在正则条件下其收敛于零。
  • 在特征数p相对于样本量n较大时,采用收缩或正则化方法以稳定偏差校正的估计。
  • 使用哮喘研究中的真实DNA甲基化数据验证该方法,与Fan & Han(2017)和Wang et al.(2017)等现有方法进行比较。

实验结果

研究问题

  • RQ1在存在未观测混杂因子的高维回归中,现有估计器的渐近偏差在多大程度上依赖于潜在因子的可估性?
  • RQ2是否可以构建一种校正估计器,即使未观测协变量仅部分可估,其渐近分布仍与OLS相同?
  • RQ3在高维数据中,未观测的细胞类型异质性在哮喘与DNA甲基化之间观察到的关联中,其中介作用有多大?
  • RQ4在真实生物数据中,该方法在错误发现率控制和统计功效方面与现有方法相比表现如何?

主要发现

  • 所提出的估计器在渐近意义上无偏,且当所有协变量均可观测时,其渐近分布与OLS相同,即使未观测因子仅弱可估也成立。
  • 研究表明,当观测协变量与未观测协变量之间的相关性非零时,Fan & Han(2017)和Wang et al.(2017)等现有方法存在渐近偏差。
  • 在对已发表DNA甲基化数据的重新分析中,该方法揭示了哮喘相关甲基化变异的大部分是由细胞组成介导的,而非直接的遗传效应。
  • 该方法纠正了先前方法未能捕捉到的哮喘与潜在细胞类型异质性之间相关性的低估问题。
  • 偏差校正项的方差为O(1/√p)量级,随着p增大而趋于零,支持了估计器的一致性。
  • 理论分析证实,即使在存在未观测混杂因子的高维设定下,该估计器仍保持与OLS相同的渐近效率。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。