Skip to main content
QUICK REVIEW

[论文解读] The Tajima heterochronous n-coalescent: inference from heterochronously sampled molecular data

Lorenzo Cappello, Amandine Véber|arXiv (Cornell University)|Apr 14, 2020
Protein Structure and Dynamics参考文献 50被引用 5
一句话总结

本文提出了一种名为Tajima异时性n-共祖的低分辨率共祖模型,通过基于采样时间而非个体标签建模谱系关系,降低了树空间复杂度,提升了计算效率,从而在从异时性采样的分子数据推断有效种群大小轨迹方面,显著优于Kingman共祖模型,在古DNA和病毒序列数据上展现出更高的可扩展性和准确性。

ABSTRACT

The observed sequence variation at a locus informs about the evolutionary history of the sample and past population size dynamics. The Kingman coalescent is used in a generative model of molecular sequence variation to infer evolutionary parameters. However, it is well understood that inference under this model does not scale well with sample size. Here, we build on recent work based on a lower resolution coalescent process, the Tajima coalescent, to model longitudinal samples. While the Kingman coalescent models the ancestry of labeled individuals, the heterochronous Tajima coalescent models the ancestry of individuals labeled by their sampling time. We propose a new inference scheme for the reconstruction of effective population size trajectories based on this model with the potential to improve computational efficiency. Modeling of longitudinal samples is necessary for applications (e.g. ancient DNA and RNA from rapidly evolving pathogens like viruses) and statistically desirable (variance reduction and parameter identifiability). We propose an efficient algorithm to calculate the likelihood and employ a Bayesian nonparametric procedure to infer the population size trajectory. We provide a new MCMC sampler to explore the space of heterochronous Tajima's genealogies and model parameters. We compare our procedure with state-of-the-art methodologies in simulations and applications.

研究动机与目标

  • 解决Kingman n-共祖模型在大规模分子数据中分析时的计算不可行性问题。
  • 提升从异时性样本中进行有效种群大小轨迹贝叶斯推断的可扩展性与混合效率。
  • 开发一种低分辨率的祖先过程,在保持关键进化信息的同时降低状态空间复杂度。
  • 在采样时间分辨率有限的古DNA和快速演化病原体(如SARS-CoV-2)上实现准确推断。
  • 提供一种计算高效且不牺牲统计准确性的Kingman共祖模型替代方案。

提出的方法

  • 提出Tajima异时性n-共祖模型,一种基于时间点上共祖事件排序而非个体标签建模谱系关系的聚合共祖过程。
  • 推导出一种快速算法以计算Tajima共祖模型下的似然,利用时间排序树拓扑结构的基数减少特性。
  • 采用基于高斯过程的贝叶斯非参数框架,对随时间变化的有效种群大小轨迹进行建模。
  • 开发一种新型MCMC采样器,高效探索异时性Tajima谱系树与模型参数的空间。
  • 在模拟与真实数据研究中,使用序贯蒙特卡洛和哈密顿蒙特卡洛作为对比基准。
  • 将该方法应用于古人类DNA与来自GISAID的SARS-CoV-2基因组数据,验证其在真实世界数据集上的性能。

实验结果

研究问题

  • RQ1像Tajima n-共祖这样的低分辨率共祖模型是否能提升从异时性数据中进行种群大小推断的计算可扩展性?
  • RQ2基于采样时间而非个体身份建模谱系关系,是否能带来更好的MCMC混合效率与更快收敛速度?
  • RQ3Tajima异时性n-共祖模型在估计有效种群大小轨迹方面,与Kingman n-共祖模型及GMRF等先进方法相比表现如何?
  • RQ4该模型在多大程度上能从古DNA与现代序列数据中恢复出已知的人口衰退或病毒爆发等人口历史事件?
  • RQ5该方法能否有效处理具有时间聚集性与有限采样分辨率的真实数据,如法国与德国的SARS-CoV-2序列?

主要发现

  • 该方法成功恢复了古人类种群的有效种群大小轨迹,显示约在26.5–19 kya开始出现下降,与已知的人口历史事件一致。
  • 对于SARS-CoV-2数据,模型估计出从2019年12月到2020年2月下旬的指数增长,峰值出现在2020年2月29日左右,随后下降,与流行病动力学一致。
  • 法国数据集的后验中位数估计与GMRF结果高度一致,两种方法均检测到相似的增长与下降模式。
  • 德国数据集显示有效种群大小估计基本恒定,可能由于2020年3月在时间和空间上的采样集中,该模型准确捕捉到了这一特征。
  • 所提出的MCMC采样器相比基于Kingman共祖模型的方法,展现出更优的混合效率与更快的收敛速度,尤其在高维树空间中表现突出。
  • 由于时间排序树拓扑结构基数减少,Tajima模型下的似然计算显著更快,使该方法能够扩展至更大规模数据集。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。