Skip to main content
QUICK REVIEW

[论文解读] High-dimensional genome-wide association study and misspecified mixed model analysis

Jiming Jiang, Cong Li|arXiv (Cornell University)|Apr 9, 2014
Genetic Associations and Epidemiology参考文献 27被引用 4
一句话总结

本文在高维线性混合模型(LMMs)模型误设的背景下,为全息最大似然(REML)估计提供了理论依据,该方法常用于全基因组关联研究(GWAS)。尽管假设所有SNP均对遗传力有贡献,而实际上仅稀疏子集真正具有影响,作者证明了方差成分的REML估计量是一致的,且收敛速率为 $O_P(\sqrt{\log n/n})$,并具有渐近正态性,支持了在实践中使用LMM进行遗传力估计的稳健性。

ABSTRACT

We study behavior of the restricted maximum likelihood (REML) estimator under a misspecified linear mixed model (LMM) that has received much attention in recent gnome-wide association studies. The asymptotic analysis establishes consistency of the REML estimator of the variance of the errors in the LMM, and convergence in probability of the REML estimator of the variance of the random effects in the LMM to a certain limit, which is equal to the true variance of the random effects multiplied by the limiting proportion of the nonzero random effects present in the LMM. The aymptotic results also establish convergence rate (in probability) of the REML estimators as well as a result regarding convergence of the asymptotic conditional variance of the REML estimator. The asymptotic results are fully supported by the results of empirical studies, which include extensive simulation studies that compare the performance of the REML estimator (under the misspecified LMM) with other existing methods.

研究动机与目标

  • 为在高维全基因组关联研究(GWAS)中广泛使用线性混合模型(LMMs)提供理论依据,尽管存在模型误设。
  • 研究当真实模型为稀疏但假设的LMM将所有SNP视为随机效应时,受限最大似然(REML)估计量的渐近行为。
  • 在样本量和随机效应数量均增长的高维渐近条件下,建立REML估计量的一致性、收敛速率及渐近方差性质。
  • 提出并形式化高维设定下误设混合模型分析(MMMA)的概念及其渐近性质。

提出的方法

  • 作者分析了在随机效应数量随样本量增长的高维LMM中,REML估计量的渐近性质。
  • 他们应用随机矩阵理论(RMT)推导了方差成分(特别是误差方差和随机效应方差)的REML估计量的极限分布。
  • 分析设定为真实模型为稀疏的(仅部分SNP具有非零效应),但假设的LMM将所有SNP视为随机效应。
  • 关键方程涉及误差方差 $\hat{\sigma}_\epsilon^2$ 和方差成分比 $\hat{\gamma}$ 的渐近行为,收敛速率通过局部鞅和迹相关论证推导。
  • 推导过程使用条件矩和方差分解,表明 $\hat{\sigma}_\epsilon^2 - \sigma_{\epsilon 0}^2 = O_P(\sqrt{\log n/n})$ 且 $\hat{\gamma} - \gamma_* = O_P(\sqrt{\log n/n})$,其中 $\gamma_*$ 为非零随机效应的极限比例。
  • 理论结果通过与其它方法比较的大量模拟研究得到支持,验证了渐近预测的准确性。

实验结果

研究问题

  • RQ1当LMM被误设(即假设所有SNP均为随机效应,而仅稀疏子集真正活跃)时,REML估计量对方差成分是否仍保持一致?
  • RQ2在高维渐近条件下,误差方差和方差成分比的REML估计量的收敛速率如何?
  • RQ3REML估计量的渐近条件方差行为如何?是否与有限样本表现一致?
  • RQ4在高维GWAS中,尽管存在模型误设,REML对遗传力的估计能否获得理论支持?
  • RQ5当随机效应数量和样本量均趋于无穷大时,REML估计量的极限行为如何?

主要发现

  • 在模型误设下,误差方差的REML估计量是一致的,随着样本量增加,依概率收敛于真实值。
  • 随机效应方差的REML估计量依概率收敛于一个极限,该极限等于真实方差乘以非零随机效应的极限比例。
  • 两个REML估计量的收敛速率均为 $O_P(\sqrt{\log n/n})$,由于高维性,该速率慢于经典 $O_P(n^{-1/2})$ 速率。
  • REML估计量的渐近条件方差依概率收敛于一个常数,支持了在高维渐近条件下推断的有效性。
  • 模拟研究证实了理论预测,表明在模型误设下,REML相较于其他方法表现稳健。
  • 结果为在GWAS中使用基于LMM的遗传力估计提供了理论支持,即使模型假设所有SNP均有贡献,也缓解了‘遗传力缺失’悖论的担忧。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。