Skip to main content
QUICK REVIEW

[论文解读] Large-scale inference of correlation among mixed-type biological traits with phylogenetic multivariate probit models

Zhenyu Zhang, Akihiko Nishimura|arXiv (Cornell University)|Dec 19, 2019
Bayesian Methods and Mixture Models参考文献 52被引用 6
一句话总结

本文提出了一种可扩展的贝叶斯框架,结合系统发育多变量 probit 模型,用于推断混合类型生物特征(连续型与二值型)之间的相关性,同时考虑共享的进化历史。通过将弹跳粒子采样器(bouncy particle sampler)与动态规划相结合,该方法实现了在高维截断正态分布中高效计算后验分布,成功分析了535株HIV分离株的24项特征,并揭示了与致病性相关的关键特征相关性。

ABSTRACT

Inferring concerted changes among biological traits along an evolutionary history remains an important yet challenging problem. Besides adjusting for spurious correlation induced from the shared history, the task also requires sufficient flexibility and computational efficiency to incorporate multiple continuous and discrete traits as data size increases. To accomplish this, we jointly model mixed-type traits by assuming latent parameters for binary outcome dimensions at the tips of an unknown tree informed by molecular sequences. This gives rise to a phylogenetic multivariate probit model. With large sample sizes, posterior computation under this model is problematic, as it requires repeated sampling from a high-dimensional truncated normal distribution. Current best practices employ multiple-try rejection sampling that suffers from slow-mixing and a computational cost that scales quadratically in sample size. We develop a new inference approach that exploits 1) the bouncy particle sampler (BPS) based on piecewise deterministic Markov processes to simultaneously sample all truncated normal dimensions, and 2) novel dynamic programming that reduces the cost of likelihood and gradient evaluations for BPS to linear in sample size. In an application with 535 HIV viruses and 24 traits that necessitates sampling from a 12,840-dimensional truncated normal, our method makes it possible to estimate the across-trait correlation and detect factors that affect the pathogen's capacity to cause disease. This inference framework is also applicable to a broader class of covariance structures beyond comparative biology.

研究动机与目标

  • 解决在调整共享系统发育历史的前提下,推断混合类型生物特征(连续型与离散型)之间协同进化变化的挑战。
  • 克服大规模系统发育模型中高维截断正态分布采样的计算瓶颈。
  • 开发一种在高维设置下,使似然与梯度计算评估的计算复杂度随样本量线性增长的方法。
  • 实现在HIV进化等复杂生物系统中对跨特征相关性的可靠推断。
  • 通过在扩散协方差结构上施加约束,解决现有多变量 probit 模型中的可识别性问题。

提出的方法

  • 提出一种系统发育多变量 probit 模型,通过在未知系统发育树上演化潜变量,联合建模连续型与二值型特征。
  • 采用潜变量框架,其中二值型特征由未观测到的连续潜变量上的阈值决定,潜变量过程在树上遵循布朗运动。
  • 应用基于分段确定性马氏过程的弹跳粒子采样器(BPS),以高效采样高维截断正态分布。
  • 提出一种新颖的动态规划算法,将似然与梯度计算的计算成本从O(N²)降低至O(N),其中N为样本量。
  • 对扩散协方差矩阵施加约束,以解决先前模型(如Cybis等,2015)中存在的可识别性问题。
  • 采用启发式方法,将BPS中的总行进时间设置为与协方差矩阵最大特征值的平方根成正比,以确保良好的混合性能。

实验结果

研究问题

  • RQ1如何在大规模数据集中,通过考虑共享的进化历史,可靠地推断混合类型生物特征(连续型与二值型)之间的相关性?
  • RQ2在系统发育模型中,有哪些计算策略可实现对高维截断正态分布似然与梯度计算的可扩展性?
  • RQ3弹跳粒子采样器能否被有效适配于系统发育多变量 probit 模型中产生的高维截断正态目标分布?
  • RQ4在大规模设置下,该方法与现有MCMC方法(如多试拒绝采样)相比,在效率与准确性方面表现如何?
  • RQ5在混合类型特征模型中,为确保模型可识别性,对扩散协方差矩阵需要施加何种约束?

主要发现

  • 所提出的方法成功从535株HIV分离株与24项混合类型特征中产生的12,840维截断正态分布中进行采样。
  • 当总行进时间 $ t_{\text{total}} = 0.01\beta\text{max}} $ 时,BPS在每小时的中位有效样本量最高,优于测试范围内其他 $ t_{\text{total}} $ 取值。
  • 动态规划方法将似然与梯度计算的计算成本从二次方降低至线性,实现了对大规模数据集的可扩展性。
  • 该方法检测到HIV致病性中的显著跨特征相关性,识别出影响病毒致病能力的关键因素。
  • 通过在扩散协方差矩阵上施加约束,解决了先前模型中的可识别性问题,提升了MCMC混合性能与参数估计效果。
  • 基于协方差矩阵最大特征值的 $ t_{\text{total}} $ 启发式设置在不同运行中均表现出稳健与高效。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。