[论文解读] Gaussian Process-Based Bayesian Nonparametric Inference of Population Trajectories from Gene Genealogies
本文提出了一种基于高斯过程的贝叶斯非参数方法,利用共祖先模型从系统发育树推断种群大小轨迹。通过将共祖先事件时间视为非齐次点过程,该方法利用先进的泊松过程非参数推断技术,在模拟和真实病毒数据分析中,相比现有的高斯马尔可夫随机场(GMRF)方法,实现了更高的准确性、精确度以及更真实的不确定性估计。
Changes in population size influence genetic diversity of the population and, as a result, leave a signature of these changes in individual genomes in the population. We are interested in the inverse problem of reconstructing past population dynamics from genomic data. We start with a standard framework based on the coalescent, a stochastic process that generates genealogies connecting randomly sampled individuals from the population of interest. These genealogies serve as a glue between the population demographic history and genomic sequences. It turns out that only the times of genealogical lineage coalescences contain information about population size dynamics. Viewing these coalescent times as a point process, estimating population size trajectories is equivalent to estimating a conditional intensity of this point process. Therefore, our inverse problem is similar to estimating an inhomogeneous Poisson process intensity function. We demonstrate how recent advances in Gaussian process-based nonparametric inference for Poisson processes can be extended to Bayesian nonparametric estimation of population size dynamics under the coalescent. We compare our Gaussian process (GP) approach to one of the state of the art Gaussian Markov random field (GMRF) methods for estimating population trajectories. Using simulated data, we demonstrate that our method has better accuracy and precision. Next, we analyze two genealogies reconstructed from real sequences of hepatitis C and human Influenza A viruses. In both cases, we recover more believed aspects of the viral demographic histories than the GMRF approach. We also find that our GP method produces more reasonable uncertainty estimates than the GMRF method.
研究动机与目标
- 开发一种灵活的非参数贝叶斯方法,用于从系统发育数据估计有效种群大小轨迹,而无需假设固定网格或分段常数函数。
- 克服现有非参数方法的局限性,这些方法依赖于任意离散化或具有固定变化点的分段连续先验。
- 通过将现代高斯过程技术适配至共祖先过程,提升估计的准确性和不确定性量化。
- 为未来扩展至多变量建模、多基因座以及与分子序列数据的整合提供支持。
- 提供一种计算上可行的框架,避免在MCMC采样过程中预先指定网格。
提出的方法
- 将具有可变种群大小的共祖先过程重新表述为非齐次点过程,其中共祖先时间被视为连续时间点过程中的事件。
- 使用均值为零的高斯过程先验对这一点过程的强度函数进行非参数建模,协方差矩阵控制平滑度。
- 开发了一种基于薄化算法的数据增广方案——借鉴Adams等人(2009)针对泊松过程的方法——专门用于共祖先过程,以实现后验计算。
- 采用非GMRF高斯过程先验(例如,布朗运动、Ornstein-Uhlenbeck过程)以通过稀疏精度矩阵确保计算可行性。
- 通过MCMC进行后验推断,有效种群大小轨迹作为高斯过程样本后验中位数进行估计。
- 该框架设计为可扩展至多变量设置,支持种群大小与外部时间序列的联合建模。
实验结果
研究问题
- RQ1基于高斯过程的非参数推断是否能相比现有GMRF方法提升种群大小轨迹估计的准确性和精确度?
- RQ2所提出的方法在种群大小动态的不确定性处理方面表现如何,特别是与GMRF方法相比?
- RQ3能否通过将共祖先时间视为具有随机强度函数的点过程,将该方法适配至共祖先模型?
- RQ4当应用于具有复杂种群历史的现实病毒系统发育树时,该方法是否保持鲁棒性和计算可行性?
- RQ5该框架未来能否扩展以整合系统发育不确定性及多个基因座?
主要发现
- 在模拟中,基于高斯过程的方法相比最先进的GMRF方法,在估计种群大小轨迹方面表现出更高的准确性和精确度。
- 该方法产生的不确定性估计比GMRF方法更合理、更真实,尤其是在共祖先事件稀疏的区域。
- 在丙型肝炎病毒和人类甲型流感病毒系统发育数据上的应用中,GP方法恢复了比GMRF方法更符合生物学实际的种群历史。
- 所选GP先验(如布朗运动、Ornstein-Uhlenbeck过程)对结果影响极小,表明对先验设定具有鲁棒性。
- 布朗运动先验中的精度参数对扰动表现出低敏感性,支持推断的稳定性。
- 该方法可为未来扩展至种群大小与外部协变量(如环境或流行病学时间序列)的多变量建模提供支持。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。