Skip to main content
QUICK REVIEW

[Paper Review] Distance-based species tree estimation: information-theoretic trade-off between number of loci and sequence length under the coalescent

Elchanan Mossel, Sébastien Roch|arXiv (Cornell University)|Apr 21, 2015
Identification and Quantification in Food42 references6 citations
TL;DR

This paper establishes an information-theoretic trade-off between the number of loci and sequence length in distance-based species tree estimation under the multispecies coalescent model. By linking the problem to sparse signal detection, it proves that detecting a species tree branch of length $ f $ requires $ m = \Theta(1/(f^2\sqrt{k})) $ loci, where $ k $ is the sequence length, revealing a fundamental limit on data requirements for accurate phylogeny reconstruction.

ABSTRACT

We consider the reconstruction of a phylogeny from multiple genes under the multispecies coalescent. We establish a connection with the sparse signal detection problem, where one seeks to distinguish between a distribution and a mixture of the distribution and a sparse signal. Using this connection, we derive an information-theoretic trade-off between the number of genes, $m$, needed for an accurate reconstruction and the sequence length, $k$, of the genes. Specifically, we show that to detect a branch of length $f$, one needs $m = Θ(1/[f^{2} \sqrt{k}])$.

Motivation & Objective

  • To understand the fundamental data requirements for accurate distance-based species tree estimation under the multispecies coalescent model.
  • To quantify the trade-off between the number of loci $ m $ and sequence length $ k $ needed to detect a species tree branch of length $ f $.
  • To establish a theoretical detection boundary for phylogenetic reconstruction by connecting it to the sparse signal detection problem.
  • To provide information-theoretic lower and upper bounds on the number of loci required for consistent species tree estimation.

Proposed method

  • Formalizes species tree estimation as a hypothesis testing problem between two distributions: one with no signal (null) and one with a sparse signal (alternative), analogous to sparse signal detection.
  • Models gene sequences under the Jukes-Cantor model and uses pairwise sequence distances as sufficient statistics for tree reconstruction.
  • Applies the Berry-Esseen theorem and concentration inequalities to bound the sampling distribution of distance quantiles across loci.
  • Develops a two-phase algorithm: first estimating a threshold $ \hat{p} $ from the $ C/\sqrt{k} $-quantile of distances, then comparing the fraction of genes below this threshold across datasets.
  • Uses a partitioned data structure to control dependencies and ensure statistical independence in hypothesis testing.
  • Derives asymptotic bounds on $ m $ using concentration and tail probability arguments, showing that $ m = \Theta(1/(f^2\sqrt{k})) $ is necessary and sufficient for detection with high probability.

Experimental results

Research questions

  • RQ1What is the minimum number of loci $ m $ required to detect a species tree branch of length $ f $, given a sequence length $ k $, under the multispecies coalescent?
  • RQ2How does the required number of loci scale with the branch length $ f $ and sequence length $ k $ in distance-based species tree estimation?
  • RQ3Can the sparse signal detection framework be used to derive information-theoretic limits for multispecies coalescent-based phylogeny reconstruction?
  • RQ4What is the fundamental detection boundary for distinguishing between two species tree topologies using multiple loci of finite length?

Key findings

  • The number of loci $ m $ required to detect a branch of length $ f $ scales as $ \Theta(1/(f^2\sqrt{k})) $, establishing a precise information-theoretic trade-off between $ m $ and $ k $.
  • The detection boundary is derived by reducing the phylogeny reconstruction problem to a sparse signal detection problem, where the signal corresponds to a branch of length $ f $.
  • The proposed two-phase algorithm achieves high-probability detection of species tree topology with $ m \geq c'/(f^2\sqrt{k}) $ loci for a sufficiently large constant $ c' $.
  • The method is robust even when $ f \ll 1/k $, as the second phase of the algorithm compares the fraction of genes below a threshold, overcoming quantization issues from finite $ k $.
  • The analysis shows that the detection boundary is sharp: if $ m $ is smaller than $ c/(f^2\sqrt{k}) $, detection fails with high probability.
  • The results apply to distance-based methods and provide a theoretical foundation for understanding data requirements in multispecies coalescent phylogenetics.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.