Skip to main content
QUICK REVIEW

[论文解读] In the Light of Deep Coalescence: Revisiting Trees Within Networks

Jiafan Zhu, Yun Yu|arXiv (Cornell University)|Jun 23, 2016
Genetic diversity and population structure参考文献 12被引用 6
一句话总结

本文提出在物种网推断中用“亲代树”(parental trees)替代传统的“显示树”(displayed trees),以更好地反映深共线性(deep coalescence)效应。通过将亲代树定义为基因树直接继承的进化历史,作者证明仅依赖显示树在不完全谱系分选(incomplete lineage sorting)情况下不足以实现准确的网络推断,并提出一种基于聚类的方法:先从基因树聚类中推断亲代树,再重建网络,从而提升准确性。

ABSTRACT

Phylogenetic networks model reticulate evolutionary histories. The last two decades have seen an increased interest in establishing mathematical results and developing computational methods for inferring and analyzing these networks. A salient concept underlying a great majority of these developments has been the notion that a network displays a set of trees and those trees can be used to infer, analyze, and study the network. In this paper, we show that in the presence of coalescence effects, the set of displayed trees is not sufficient to capture the network. We formally define the set of parental trees of a network and make three contributions based on this definition. First, we extend the notion of anomaly zone to phylogenetic networks and report on anomaly results for different networks. Second, we demonstrate how coalescence events could negatively affect the ability to infer a species tree that could be augmented into the correct network. Third, we demonstrate how a phylogenetic network can be viewed as a mixture model that lends itself to a novel inference approach via gene tree clustering. Our results demonstrate the limitations of focusing on the set of trees displayed by a network when analyzing and inferring the network. Our findings can form the basis for achieving higher accuracy when inferring phylogenetic networks and open up new venues for research in this area, including new problem formulations based on the notion of a network's parental trees.

研究动机与目标

  • 解决在存在深共线性(不完全谱系分选)时,仅使用显示树进行物种网推断的局限性。
  • 定义并形式化亲代树的概念,作为网络内基因树进化更准确的表示。
  • 将异常区(anomaly zone)概念扩展至网络,表明最可能的基因树可能不在任何显示树中。
  • 提出一种基于聚类基因树并重建亲代树、随后构建网络的新推断方法。
  • 证明现有骨干树推断方法在广泛深共线性条件下失效,强调需要针对具有种群效应的网状进化设计新方法。

提出的方法

  • 将物种网的亲代树集合定义为与网中节点和分支直接对应的进化路径所生成的树。
  • 正式定义网络的异常区,识别在分支长度和遗传概率区域中,最可能的基因树并非亲代树的情况。
  • 提出基于聚类的推断流程:使用相异度度量对基因树进行聚类,利用最小深共线性(MDC)方法从每类中推断亲代树,并将这些亲代树合并为网络。
  • 在性能评估中使用Robinson-Foulds(RF)距离度量推断亲代树与真实亲代树之间的相似性。
  • 在PhyloNet中实现该方法,采用最小权边覆盖算法匹配推断树与真实树。
  • 在模拟数据上评估性能,使用不同数量的基因树(50–250),测量亲代树恢复的错误率。

实验结果

研究问题

  • RQ1能否将异常区概念从物种树有意义地扩展至物种网,基于完整的亲代树集合而非单一指定的物种树?
  • RQ2深共线性在多大程度上导致最可能的基因树偏离网络中所有亲代树,从而破坏标准推断方法的可靠性?
  • RQ3基于聚类的方法,即从基因树组中推断亲代树,是否能优于传统方法中试图使网络显示所有输入基因树的做法?
  • RQ4为何现有骨干树推断方法在广泛深共线性条件下失效?这对网络推断有何启示?
  • RQ5将物种网建模为亲代树的混合体而非显示树的集合,是否具有根本性优势?

主要发现

  • 当深共线性发生时,网络所显示的树集合不足以捕捉基因树的进化过程,因为最可能的基因树可能不在其中。
  • 网络的异常区基于完整的亲代树集合进行定义,且研究证明此类区域确实存在,并可能影响推断准确性。
  • 现有在网中推断骨干树或物种树的方法在广泛深共线性条件下表现不佳,表明亟需新方法。
  • 所提出的基于聚类的方法在使用250个或更多基因树时,亲代树恢复错误率约为2%;在使用50个基因树时,错误率约为10%。
  • 在模拟中,若网络需显示所有输入基因树,其结果与真实网络存在显著差异,凸显了传统方法的缺陷。
  • 亲代树为网络推断提供了更准确的基础,且在聚类基因树后推断亲代树的方法为物种网重建开辟了有前景的新路径。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。