[论文解读] Modeling population structure under hierarchical Dirichlet processes
本文提出一种基于层次狄利克雷过程(HDP)的贝叶斯非参数模型,用于在考虑基因位点间连锁不平衡(LD)的前提下推断群体混合。通过建模相邻遗传标记间的相关祖先状态,该方法在传统方法(如STRUCTURE)的基础上实现了改进,允许祖先片段长度可变,并对祖先群体数量进行非参数推断,展示了在真实数据中准确检测稀有单倍型的能力。
We propose a Bayesian nonparametric model to infer population admixture, extending the Hierarchical Dirichlet Process to allow for correlation between loci due to Linkage Disequilibrium. Given multilocus genotype data from a sample of individuals, the model allows inferring classifying individuals as unadmixed or admixed, inferring the number of subpopulations ancestral to an admixed population and the population of origin of chromosomal regions. Our model does not assume any specific mutation process and can be applied to most of the commonly used genetic markers. We present a MCMC algorithm to perform posterior inference from the model and discuss methods to summarise the MCMC output for the analysis of population admixture. We demonstrate the performance of the proposed model in simulations and in a real application, using genetic data from the EDAR gene, which is considered to be ancestry-informative due to well-known variations in allele frequency as well as phenotypic effects across ancestry. The structure analysis of this dataset leads to the identification of a rare haplotype in Europeans.
研究动机与目标
- 解决现有混合模型在处理紧密相邻基因位点间连锁不平衡(LD)方面的局限性。
- 开发一种方法,实现对祖先群体数量的非参数推断,无需预设固定的K值。
- 实现对染色体区段到祖先群体的准确分配,反映混合基因组的镶嵌特性。
- 在不假设特定突变过程的前提下建模群体结构,使其适用于多种遗传标记。
- 提供一个完整的贝叶斯框架,包含MCMC推断和后验总结工具,用于群体混合分析。
提出的方法
- 将层次狄利克雷过程(HDP)扩展用于建模位点间相关的祖先状态,通过中国人餐厅过程(CRP)表示群体分配,以捕捉连锁不平衡。
- 使用吉布斯采样器,对群体成员关系(z_il)、祖先比例(n_ik, m_ik)、等位基因频率(θ_kl)以及超参数(α, α₀, r, μ_l)进行条件更新。
- 对SNP数据采用Beta-Binomial共轭结构,其中θ_kl在给定各群体中观察到的基因型条件下服从Beta分布。
- 通过辅助变量和Polya-Gamma数据增广方法,实现对浓度参数α和α₀的高效后验更新。
- 对基测度H中的泊松过程速率r和超参数μ_l采用随机游走Metropolis步骤进行更新。
- 对基测度H采用非参数先验,以灵活建模祖先等位基因频率,而无需假设固定的参数形式。
实验结果
研究问题
- RQ1当位点因连锁不平衡而相关时,如何对群体混合进行建模?
- RQ2贝叶斯非参数模型是否可以在不预先指定K值的情况下推断祖先群体数量?
- RQ3考虑连锁不平衡在多大程度上能提高染色体祖先推断的准确性?
- RQ4如何设计MCMC采样方法,以高效探索混合模型中复杂且高维的后验分布?
- RQ5利用该模型能对稀有遗传变异(如EDAR单倍型)获得哪些新见解?
主要发现
- 该模型成功检测到欧洲人群中EDAR基因内的一个稀有单倍型,该单倍型此前被标准方法所忽略。
- 对祖先群体数量的后验推断为非参数确定,模型能自动适应数据驱动的聚类数量。
- 引入LD建模后,对祖先片段长度和染色体区域祖先来源的推断更加准确。
- 模拟结果表明,该模型在LD条件下优于标准HDP和STRUCTURE,能更准确恢复真实的群体结构。
- MCMC算法收敛良好,并能为祖先比例和群体分配提供可靠的后验总结。
- 该方法在不同遗传标记下均表现稳健,且无需对潜在突变过程做任何假设。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。