Skip to main content
QUICK REVIEW

[论文解读] A model selection approach to genome wide association studies

Florian Frommlet, Felix Ruhaltinger|arXiv (Cornell University)|Oct 1, 2010
Genetic Associations and Epidemiology参考文献 39被引用 9
一句话总结

本文提出一种基于改进贝叶斯信息准则(mBIC 和 mBIC2)的模型选择方法,用于全基因组关联研究(GWAS),以提高在复杂性状中检测致病SNP的能力。通过同时建模多个SNP而非单独测试,该方法提升了统计功效,并减少了因SNP之间虚假相关性导致的假阳性结果,在模拟实验和HapMap基因表达数据的真实数据分析中,其表现优于传统多重检验方法。

ABSTRACT

For the vast majority of genome wide association studies (GWAS) published so far, statistical analysis was performed by testing markers individually. In this article we present some elementary statistical considerations which clearly show that in case of complex traits the approach based on multiple regression or generalized linear models is preferable to multiple testing. We introduce a model selection approach to GWAS based on modifications of Bayesian Information Criterion (BIC) and develop some simple search strategies to deal with the huge number of potential models. Comprehensive simulations based on real SNP data confirm that model selection has larger power than multiple testing to detect causal SNPs in complex models. On the other hand multiple testing has substantial problems with proper ranking of causal SNPs and tends to detect a certain number of false positive SNPs, which are not linked to any of the causal mutations. We show that this behavior is typical in GWAS for complex traits and can be explained by an aggregated influence of many small random sample correlations between genotypes of a SNP under investigation and other causal SNPs. We believe that our findings at least partially explain problems with low power and nonreplicability of results in many real data GWAS. Finally, we discuss the advantages of our model selection approach in the context of real data analysis, where we consider publicly available gene expression data as traits for individuals from the HapMap project.

研究动机与目标

  • 为解决单标记检验在GWAS中的局限性,特别是复杂遗传结构下统计功效低和假阳性率高的问题。
  • 开发一种基于改进BIC准则(mBIC和mBIC2)的模型选择框架,适用于具有少量真实信号的高维SNP数据。
  • 通过模拟实验和真实数据验证,证明模型选择在检测致病SNP及准确排序方面优于多重检验方法。
  • 解释单标记检验中p值排序不稳定的成因,即致病SNP与非致病SNP之间随机相关性的影响。
  • 在HapMap基因表达数据中,识别出传统多重检验方法未能检测到的新型反式作用SNP关联。

提出的方法

  • 采用多元回归或广义线性模型框架,联合建模多个SNP,避免单变量检验的局限性。
  • 使用mBIC和mBIC2作为模型选择准则,其惩罚项针对高维设置和稀疏信号检测进行了优化。
  • 采用两阶段搜索策略:首先通过前向选择或筛选方法识别有前景的SNP组合,再通过局部搜索或逐步法进行优化。
  • 利用mBIC2在更广泛的稀疏条件下具有渐近最优性,提升模型选择的一致性。
  • 利用HapMap项目的真实SNP数据模拟GWAS数据,验证模型在真实遗传相关结构下的性能表现。
  • 将模型选择结果与标准多重检验程序的结果进行比较,重点关注统计功效、假发现率以及真实致病SNP的检测能力。

实验结果

研究问题

  • RQ1在复杂遗传模型中,使用mBIC或mBIC2的模型选择是否在检测致病SNP方面显著优于单标记检验?
  • RQ2非致病SNP与致病SNP之间的随机相关性在多大程度上扭曲了单标记GWAS中的p值排序?
  • RQ3模型选择方法能否检测到多重检验方法遗漏的生物学上相关的反式作用SNP关联?
  • RQ4在高维、稀疏的GWAS设置下,mBIC和mBIC2在模型选择一致性和假阳性控制方面表现如何?
  • RQ5样本量对模型选择在真实GWAS数据中识别复杂多SNP关联能力有何影响?

主要发现

  • 使用mBIC和mBIC2的模型选择在检测致病SNP方面显著优于多重检验,尤其在存在多个相互作用位点的复杂模型中表现更优。
  • 单标记检验因未建模的致病变异导致残差平方和被放大,从而造成统计功效低下,这是多重检验校正无法解决的关键局限。
  • 非致病SNP与致病SNP之间的微小随机相关性会导致虚假的p值排序,引发假阳性结果并导致真实致病变异无法被检测到。
  • 在HapMap基因表达特征的真实数据分析中,mBIC2检测到原始多重检验分析中未发现的新型反式SNP,包括原始研究遗漏的一个顺式SNP。
  • 该模型选择方法识别出多个基因共享的调控区域,例如染色体6上(32.5–32.8 Mb)存在强信号,多个反式SNP与致病变异高度相关。
  • 尽管样本量较小限制了模型复杂度(最大模型大小为11个SNP),该方法仍检测到具有生物学合理性的关联,表明在更大队列中具有更大潜力。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。