Skip to main content
QUICK REVIEW

[论文解读] Improving heritability estimation by a variable selection approach in sparse high dimensional linear mixed models

Anna Bonnet, Céline Lévy‐Leduc|arXiv (Cornell University)|Jul 22, 2015
Genetic Associations and Epidemiology参考文献 27被引用 7
一句话总结

本文提出 EstHer,一种针对稀疏高维线性混合模型的新型遗传力估计方法,通过在最大似然估计前先利用变量选择识别致病遗传变异,从而提高估计精度。该方法在高稀疏性条件下显著缩小了置信区间宽度,同时保持较低计算成本,已在模拟数据和神经解剖学 Imagen 数据上得到验证。

ABSTRACT

Motivated by applications in neuroanatomy, we propose a novel methodology for estimating the heritability which corresponds to the proportion of phenotypic variance which can be explained by genetic factors. Estimating this quantity for neuroanatomical features is a fundamental challenge in psychiatric disease research. Since the phenotypic variations may only be due to a small fraction of the available genetic information, we propose an estimator of the heritability that can be used in high dimensional sparse linear mixed models. Our method consists of three steps. Firstly, a variable selection stage is performed in order to recover the support of the genetic effects -- also called causal variants -- that is to find the genetic effects which really explain the phenotypic variations. Secondly, we propose a maximum likelihood strategy for estimating the heritability which only takes into account the causal genetic effects found in the first step. Thirdly, we compute the standard error and the 95% confidence interval associated to our heritability estimator thanks to a nonparametric bootsrap approach. Our main contribution consists in providing an estimation of the heritability with standard errors substantially smaller than methods without variable selection when the genetic effects are very sparse. Since the real genetic architecture is in general unknown in practice, we also propose an empirical criterion which allows the user to decide whether it is relevant to apply a variable selection based approach or not. We illustrate the performance of our methodology on synthetic and real neuroanatomic data coming from the Imagen project. We also show that our approach has a very low computational burden and is very efficient from a statistical point of view.

研究动机与目标

  • 解决在仅少数 SNP 为致病因素的神经解剖学特征中估计遗传力的挑战。
  • 克服标准方法假设所有 SNP 均等贡献的局限性,避免导致遗传力估计不精确。
  • 开发一种计算高效的算法,利用遗传效应的稀疏性以提高统计精度。
  • 提供一个决策准则,指导用户在基于变量选择(如 EstHer)与非选择方法(如 HiLMM)之间进行选择。
  • 实现对贡献于表型方差的生物相关 SNP 的识别,以支持后续功能分析。

提出的方法

  • 在高维数据中应用变量选择程序(如 Lasso 或类似方法),以识别致病遗传效应的支持集(即非零 SNP 效应)
  • 在仅包含所选致病变异的简化模型上,使用最大似然法估计遗传力
  • 通过非参数自 resampling(自助法)计算标准误和 95% 置信区间,以确保推断稳健性
  • 基于合成数据上估计误差最小化的原则,采用数据驱动方法校准变量选择的阈值
  • 提出一个基于不同阈值下置信区间重叠程度的实证准则,以评估变量选择是否提升了估计的稳定性
  • 在 R 包 EstHer 中实现完整分析流程,该包可在 CRAN 及第一作者个人网页上获取

实验结果

研究问题

  • RQ1在具有稀疏遗传效应的高维线性混合模型中,变量选择能否提高遗传力估计的精度?
  • RQ2与标准遗传力估计器(如 HiLMM、GCTA)相比,所提出方法在置信区间宽度和准确性方面表现如何?
  • RQ3在真实神经解剖学数据中,变量选择的最佳阈值是什么?该阈值能否在不了解真实遗传结构的前提下进行校准?
  • RQ4在何种条件下,基于变量选择的方法在遗传力估计中相比标准方法具有统计优势?
  • RQ5所提出方法能否识别出对表型变异具有生物学意义的 SNP 集合,从而支持后续功能分析?

主要发现

  • 与 HiLMM 和 GCTA 相比,EstHer 在遗传力估计中实现了显著更窄的 95% 置信区间,尤其在致病 SNP 比例较低(高稀疏性)时表现更优。
  • 对于 pa、amy 和 acc 等表型,EstHer 产生的置信区间不仅更窄,且完全包含于 HiLMM 和 GCTA 的置信区间之内,表明其精度更高。
  • 基于模拟实验中真实遗传力与估计遗传力绝对差值最小化,Imagene 数据中变量选择的最优阈值为 0.79。
  • 基于不同阈值下置信区间重叠数量的实证决策准则,有效识别出 EstHer 显著优于非选择方法的表型。
  • 该方法计算负担低,适用于大规模遗传数据集,且已实现为开源 R 包 EstHer。
  • EstHer 提供候选致病 SNP 列表,可提供超越遗传力估计本身的生物学见解,而这是标准 LMM 方法所不具备的。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。