Skip to main content
QUICK REVIEW

[Paper Review] The phylogenetic Kantorovich-Rubinstein metric for environmental sequence samples

Steven N. Evans, Frederick A. Matsen|arXiv (Cornell University)|May 10, 2010
Gene expression and cancer classification4 citations
TL;DR

This paper establishes that the weighted UniFrac distance between environmental microbial sequence samples is equivalent to the phylogenetic Kantorovich-Rubinstein (KR) metric, a well-known optimal transport distance on a phylogenetic tree. It demonstrates that this metric can be computed as an integral over the tree, generalizes to L^p Zolotarev-type distances, and provides a computable approximation for permutation test p-values using a Gaussian process functional, with the L^2 case linked to a linear combination of chi-squared variables.

ABSTRACT

Using modern technology, it is now common to survey microbial communities by sequencing DNA or RNA extracted in bulk from a given environment. Comparative methods are needed that indicate the extent to which two communities differ given data sets of this type. UniFrac, a method built around a somewhat ad hoc phylogenetics-based distance between two communities, is one of the most commonly used tools for these analyses. We provide a foundation for such methods by establishing that if one equates a metagenomic sample with its empirical distribution on a reference phylogenetic tree, then the weighted UniFrac distance between two samples is just the classical Kantorovich-Rubinstein (KR) distance between the corresponding empirical distributions. We demonstrate that this KR distance and extensions of it that arise from incorporating uncertainty in the location of sample points can be written as a readily computable integral over the tree, we develop $L^p$ Zolotarev-type generalizations of the metric, and we show how the p-value of the resulting natural permutation test of the null hypothesis "no difference between the two communities" can be approximated using a functional of a Gaussian process indexed by the tree. We relate the $L^2$ case to an ANOVA-type decomposition and find that the distribution of its associated Gaussian functional is that of a computable linear combination of independent $\\chi_1^2$ random variables.

Motivation & Objective

  • To provide a rigorous mathematical foundation for UniFrac, a widely used method in microbial community comparison.
  • To show that weighted UniFrac is equivalent to the classical Kantorovich-Rubinstein (KR) distance between empirical distributions on a phylogenetic tree.
  • To develop computable generalizations of the KR metric using L^p Zolotarev-type distances for improved statistical inference.
  • To enable accurate p-value approximation for permutation tests of community differences using a Gaussian process indexed by the tree.
  • To connect the L^2 case of the metric to an ANOVA-type decomposition and derive the distribution of its test statistic as a linear combination of chi-squared random variables.

Proposed method

  • Represent each metagenomic sample as an empirical probability distribution on a reference phylogenetic tree.
  • Define the weighted UniFrac distance as the KR distance between two such empirical distributions, which is shown to be equivalent to the standard UniFrac formula.
  • Express the KR distance as a computable double integral over the tree: $ Z_2^2(P,Q) = \frac{1}{2}\frac{(m+n)^2}{mn} \left[ \int_T \int_T d(v,w) R(dv)R(dw) - \left( \frac{m}{m+n} \int_T \int_T d(v,w) P(dv)P(dw) + \frac{n}{m+n} \int_T \int_T d(v,w) Q(dv)Q(dw) \right) \right] $.
  • Generalize the KR metric to L^p Zolotarev-type distances to allow for robustness and flexibility in different statistical settings.
  • Approximate the null distribution of the test statistic using a functional of a Gaussian process indexed by the tree, enabling permutation test p-value estimation.
  • Derive the exact distribution of the L^2 KR distance statistic as a linear combination of independent $ \chi_1^2 $ random variables, enabling efficient inference.

Experimental results

Research questions

  • RQ1Is there a deeper mathematical justification for the weighted UniFrac metric beyond its ad hoc definition?
  • RQ2Can the weighted UniFrac distance be interpreted as a known metric in optimal transport theory?
  • RQ3How can the KR metric be efficiently computed and generalized for statistical inference in microbial community analysis?
  • RQ4What is the limiting distribution of the test statistic under the null hypothesis of no difference between communities?
  • RQ5Can the L^2 KR distance be decomposed in a way analogous to ANOVA, and what is the distribution of its components?

Key findings

  • The weighted UniFrac distance is mathematically equivalent to the Kantorovich-Rubinstein distance between empirical probability measures on a phylogenetic tree.
  • The KR distance can be computed as a double integral over the tree, providing a computationally feasible method for large-scale microbial community comparisons.
  • The L^p Zolotarev-type generalizations of the KR metric are well-defined and can be used to extend the framework to robust statistical inference.
  • The p-value for a permutation test of community difference can be approximated using a functional of a Gaussian process indexed by the tree, enabling accurate inference without full resampling.
  • In the L^2 case, the test statistic's distribution is a computable linear combination of independent $ \chi_1^2 $ random variables, allowing for exact p-value computation.
  • The KR distance admits an ANOVA-type decomposition, and the between-group variation is captured by the integral expression involving the pooled distribution R.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.