Skip to main content
QUICK REVIEW

[论文解读] Detection of correlation between genotypes and environmental variables. A fast computational approach for genomewide studies

G. Guillot|arXiv (Cornell University)|Jun 5, 2012
Genetic and phenotypic traits in livestock参考文献 8被引用 6
一句话总结

该论文提出了一种快速、无需MCMC的时空广义线性混合模型(SGLMM),采用INLA-SPDE框架,用于检测等位基因频率与环境变量之间的相关性,同时考虑了种群的空间结构。与传统的Mantel检验和逻辑回归方法相比,该方法通过显式建模空间依赖性,显著提升了全基因组研究中对受选择位点的检测能力。

ABSTRACT

Genomic regions (or loci) displaying outstanding correlation with some environmental variables are likely to be under selection and this is the rationale of recent methods of identifying selected loci and retrieving functional information about them. To be efficient, such methods need to be able to disentangle the potential effect of environmental variables from the confounding effect of population history. For the routine analysis of genome-wide datasets, one also needs fast inference and model selection algorithms. We propose a method based on an explicit spatial model which is an instance of spatial generalized linear mixed model (SGLMM). For inference, we make use of the INLA-SPDE theoretical and computational framework developed by Rue et al. (2009) and Lindgren et al (2011). The method we propose allows one to quantify the correlation between genotypes and environmental variables. It works for the most common types of genetic markers, obtained either at the individual or at the population level. Analyzing simulated data produced under a geostatistical model then under an explicit model of selection, we show that the method is efficient. We also re-analyze a dataset relative to nineteen pine weevils (Hylobius abietis}) populations across Europe. The method proposed appears also as a statistically sound alternative to the Mantel tests for testing the association between genetic and environmental variables.

研究动机与目标

  • 开发一种计算高效的计算方法,用于在大规模基因组数据集中检测遗传标记与环境变量之间的相关性。
  • 在选择扫描中考虑种群的空间结构以及种群历史的混杂效应。
  • 为Mantel检验提供一种统计上稳健的替代方法,因为后者在存在空间相关性时已被证明无效。
  • 实现在无需MCMC的情况下进行模型选择和贝叶斯推断,确保适用于全基因组分析的可扩展性。
  • 适用于个体水平和种群水平的基因型数据,包括显性(AFLP)和共显性(SNP)标记。

提出的方法

  • 该方法采用空间广义线性混合模型(SGLMM),其中空间随机效应通过在三角剖分网格上的高斯马尔可夫随机场(GMRF)建模。
  • 空间依赖性通过马尔可夫(Matérn)协方差函数捕捉,并通过SPDE方法近似以提高计算效率。
  • 推断通过集成嵌套拉普拉斯近似(INLA)完成,避免了计算密集型的MCMC抽样。
  • 模型将环境变量作为固定效应,支持定量和分类预测变量。
  • 通过贝叶斯因子进行模型比较,以评估每个位点的环境关联证据。
  • 该方法适用于平面(R²)和球面(S²)空间域,适合大陆尺度的研究。

实验结果

研究问题

  • RQ1与非空间方法(如逻辑回归)相比,显式考虑空间结构的模型是否能提升对受选择位点的检测能力?
  • RQ2在基因型-环境关联研究中,考虑种群空间结构如何影响统计功效和第一类错误率?
  • RQ3在存在空间自相关的情况下,所提出的SGLMM相较于Mantel检验的性能提升程度如何?
  • RQ4在已知选择与扩散模式的模拟数据中,该方法对真实选择信号的恢复能力如何?
  • RQ5在真实基因组数据(如松树大豹象数据集)中,该方法能否可靠地识别出受选择的候选位点?

主要发现

  • 在地理统计模型下的模拟中,SGLMM正确恢复了遗传变异的空间模式,并检测到了真实的环境相关性。
  • 在符合生物学现实的选择模型下,该方法检测到了100%的20个受选择位点(索引101–120),具有高后验概率和强贝叶斯因子。
  • 在松树大豹象数据集中,SGLMM仅识别出3个显著位点(BF > 3)与昼夜温差相关,而普通逻辑回归检测到11个,表明假阳性率更低。
  • 对于其他环境变量(霜冻日数、降水量、风速),SGLMM未发现显著关联(0个位点),表明对第一类错误的控制更优。
  • SGLMM与先前分析(Joost et al., 2008)在识别最强关联位点方面高度一致,验证了该方法的可靠性。
  • 该方法展示了计算效率和可扩展性,使其适用于无需MCMC的全基因组数据集的常规分析。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。