Skip to main content
QUICK REVIEW

[论文解读] Coalescent-based species tree estimation: a stochastic Farris transform

Gautam Dasarathy, Elchanan Mossel|arXiv (Cornell University)|Jul 13, 2017
Genomics and Phylogenetic Studies参考文献 43被引用 3
一句话总结

本文提出一种随机Farris变换,使在无需分子钟假设的前提下,也能在多物种谱系分叉模型下实现基于凝聚的物种树估计。通过将非分子钟情形转化为分子钟情形,该方法在数据需求上(对数因子范围内)与先前工作保持一致,并提出了一个关键可识别性结果:即使在没有分子钟的情况下,也能从未根加权基因树分布中恢复出有根物种树。

ABSTRACT

The reconstruction of a species phylogeny from genomic data faces two significant hurdles: 1) the trees describing the evolution of each individual gene--i.e., the gene trees--may differ from the species phylogeny and 2) the molecular sequences corresponding to each gene often provide limited information about the gene trees themselves. In this paper we consider an approach to species tree reconstruction that addresses both these hurdles. Specifically, we propose an algorithm for phylogeny reconstruction under the multispecies coalescent model with a standard model of site substitution. The multispecies coalescent is commonly used to model gene tree discordance due to incomplete lineage sorting, a well-studied population-genetic effect. In previous work, an information-theoretic trade-off was derived in this context between the number of loci, $m$, needed for an accurate reconstruction and the length of the locus sequences, $k$. It was shown that to reconstruct an internal branch of length $f$, one needs $m$ to be of the order of $1/[f^{2} \sqrt{k}]$. That previous result was obtained under the molecular clock assumption, i.e., under the assumption that mutation rates (as well as population sizes) are constant across the species phylogeny. Here we generalize this result beyond the restrictive molecular clock assumption, and obtain a new reconstruction algorithm that has the same data requirement (up to log factors). Our main contribution is a novel reduction to the molecular clock case under the multispecies coalescent. As a corollary, we also obtain a new identifiability result of independent interest: for any species tree with $n \geq 3$ species, the rooted species tree can be identified from the distribution of its unrooted weighted gene trees even in the absence of a molecular clock.

研究动机与目标

  • 解决由于不完全谱系分选导致基因树与物种树不一致时重建物种系统发育关系的挑战。
  • 开发一种不依赖于严格分子钟假设的方法,以克服先前基于凝聚的方法的限制。
  • 在准确的物种树重建中保持最优数据需求——具体而言,$ m \sim 1/(f^2\sqrt{k}) $,其中 $ m $ 为基因座数量,$ k $ 为序列长度。
  • 建立新的可识别性结果:即使在没有分子钟的情况下,也能从未根加权基因树分布中恢复出有根物种树。

提出的方法

  • 作者提出一种新颖的随机Farris变换,将多物种谱系分叉模型下的基因树距离映射为适用于基于分子钟推断的形式。
  • 通过基因树距离的变换,将非分子钟情形转化为分子钟情形,利用多物种谱系分叉模型在物种树上生成基因树分布的特性。
  • 使用基于分位数的检验来推断物种关系的三元组,依赖于从序列数据中估计出的基因树距离。
  • 应用集中不等式,特别是Bernstein不等式,以控制距离估计和三元组推断中的误差概率。
  • 该方法避免了直接的基因树重建,而是直接从序列数据推断物种树拓扑结构,从而最小化基因树估计误差的传播。
  • 理论分析表明,数据需求的量级为 $ m \sim 1/(f^2\sqrt{k}) $,与分子钟假设下的先前结果一致,仅在对数因子范围内存在差异。

实验结果

研究问题

  • RQ1是否可以在不假设分子钟的前提下,于多物种谱系分叉模型下实现物种树估计,同时保持最优数据需求?
  • RQ2是否可以通过一种保持统计一致性的变换,将非分子钟情形简化为分子钟情形?
  • RQ3当不存在分子钟时,是否能够从未根加权基因树分布中识别出有根物种树?
  • RQ4在一般凝聚模型下,准确的物种树重建中,基因座数量 $ m $ 与序列长度 $ k $ 之间的信息论权衡关系为何?

主要发现

  • 所提出的随机Farris变换使得在多物种谱系分叉模型下无需分子钟假设即可实现物种树估计,且数据需求与分子钟情形下的结果保持一致。
  • 该方法实现了数据需求 $ m = \Theta(1/(f^2\sqrt{k})) $,其中 $ f $ 为内部枝长度,仅在对数因子范围内存在差异。
  • 建立了新的可识别性结果:对于任意 $ n \geq 3 $,即使在没有分子钟的情况下,也能从未根加权基因树分布中识别出有根物种树。
  • 三元组推断中的误差概率被限制在 $ \mathcal{O}(\bar{\mathbb{P}}[\mathcal{E}_{12|3}] + \phi + \alpha) $ 以内,其中 $ \mathcal{E}_{12|3} $ 表示内部枝过短的事件,$ \phi, \alpha $ 为距离估计中的误差项。
  • 通过集中不等式表明,在距离估计器的间隙与方差满足适当边界条件下,错误三元组推断的概率随基因座数量呈指数衰减。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。