Skip to main content
QUICK REVIEW

[论文解读] Identifiability of the unrooted species tree topology under the coalescent model with time-reversible substitution processes

Julia Chifman, Laura Kubatko|arXiv (Cornell University)|Jun 18, 2014
Genomics and Phylogenetic Studies参考文献 32被引用 10
一句话总结

本文在共祖先模型下,针对时间可逆的替换过程,证明了从多基因DNA序列数据中可普遍识别出物种树的无根拓扑结构。作者建立了形式化的可识别性,确认当基因树在多物种共祖先模型与可逆替换模型下生成时,物种树拓扑结构可从基因树数据中一致地推断出来。

ABSTRACT

The inference of the evolutionary history of a collection of organisms is a problem of fundamental importance in evolutionary biology. The abundance of DNA sequence data arising from genome sequencing projects has led to significant challenges in the inference of these phylogenetic relationships. Among these challenges is the inference of the evolutionary history of a collection of species based on sequence information from several distinct genes sampled throughout the genome. It is widely accepted that each individual gene has its own phylogeny, which may not agree with the species tree. Many possible causes of this gene tree incongruence are known. The best studied is incomplete lineage sorting, which is commonly modeled by the coalescent process. Numerous methods based on the coalescent process have been proposed for estimation of the phylogenetic species tree given multi-locus DNA sequence data. However, use of these methods assumes that the phylogenetic species tree can be identified from DNA sequence data at the leaves of the tree, although this has not been formally established. We prove that the unrooted topology of the $n$-leaf phylogenetic species tree is generically identifiable given observed data at the leaves of the tree that are assumed to have arisen from the coalescent process with time-reversible substitution.

研究动机与目标

  • 正式确立在多物种共祖先模型下,能否从多基因DNA序列数据中识别出无根物种树拓扑结构。
  • 解决系统发育方法中长期存在的假设:物种树拓扑结构可从基因树数据中识别。
  • 通过证明在时间可逆替换过程下的可识别性,为多基因物种树估计提供理论基础。
  • 解决由于不完全谱系分选导致的基因树不一致对物种树推断带来的挑战。

提出的方法

  • 作者采用基于多物种共祖先模型的概率框架,以模拟物种树上基因树的演化。
  • 假设采用时间可逆的核苷酸替换过程(如GTR模型),以模拟基因树上的序列演化。
  • 分析聚焦于在共祖先过程中,物种树叶节点上观察到的DNA序列的联合分布。
  • 通过证明不同无根物种树拓扑结构会产生不同的观察序列数据概率分布,从而确立普遍可识别性。
  • 该证明依赖代数几何与矩量方法,以表明物种树拓扑结构可从数据分布中恢复。

实验结果

研究问题

  • RQ1在多物种共祖先模型下,能否从多基因DNA序列数据中唯一确定无根物种树拓扑结构?
  • RQ2当基因树在不完全谱系分选和时间可逆替换过程中生成时,物种树拓扑结构是否可识别?
  • RQ3在共祖先模型下,对于具有n片叶的无根物种树,普遍可识别性假设是否成立?
  • RQ4在不假设已知根的情况下,能否从未观察到的序列数据分布中恢复物种树拓扑结构?

主要发现

  • 在多物种共祖先模型与时间可逆替换过程中,无根物种树拓扑结构可从多基因DNA序列数据中普遍识别。
  • 不同的无根物种树拓扑结构会产生不同的观察序列数据概率分布,从而实现物种树的唯一恢复。
  • 该可识别性结果在一般条件下成立,即适用于模型中的几乎所有参数值。
  • 该证明确认了现有多基因物种树推断方法的理论有效性,这些方法均基于可识别性假设。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。