[Paper Review] Limit theorems for sequences of random trees
This paper establishes laws of large numbers and an invariance principle for sequences of i.i.d. random trees in a compact metric space of rooted trees, defining a d-mean tree as the minimizer of expected distance. It proposes a Kolmogorov-Smirnov-type goodness-of-fit test based on the supremum of the empirical process, with a computationally efficient implementation via minimal cut algorithms in a derived network, enabling statistical inference on tree distributions even when mean trees are identical.
We consider a random tree and introduce a metric in the space of trees to define the ``mean tree'' as the tree minimizing the average distance to the random tree. When the resulting metric space is compact we have laws of large numbers and central limit theorems for sequence of independent identically distributed random trees. As application we propose tests to check if two samples of random trees have the same law.
Motivation & Objective
- To develop statistical tools for sequences of random trees by defining a d-mean tree as the minimizer of average distance in a metric space.
- To establish strong laws of large numbers and an invariance principle for i.i.d. random trees in compact metric spaces.
- To propose a universal goodness-of-fit test for tree distributions using the supremum of the empirical process.
- To enable practical computation of the test statistic via a transformation to a minimal cut problem in a network.
- To validate the method on simulated Galton-Watson trees and real FGF protein family data.
Proposed method
- Define a metric space of rooted trees using the product topology on {0,1}^Ṽ, where Ṽ is the set of all finite sequences indexing vertices.
- Introduce the d-mean as the tree minimizing the ν-average d-distance to a random tree with law ν.
- Prove the empirical d-mean converges almost surely to the true d-mean under compactness and uniqueness.
- Establish an invariance principle for the process (gₙ(y) − g(y)), y ∈ T, using the majorizing measure condition and Ledoux-Talagrand theorem.
- Construct a test statistic as the supremum of |gₙ(y) − g(y)| over y ∈ T, which measures deviation from the null distribution.
- Transform the computation of the supremum into a minimal cut problem in a network, enabling efficient computation via graph algorithms.
Experimental results
Research questions
- RQ1Can laws of large numbers and invariance principles be established for i.i.d. sequences of random trees in a metric space of trees?
- RQ2Does the d-mean provide a consistent estimator for the underlying distribution of random trees?
- RQ3Can the supremum of the empirical process over the tree space serve as a valid test statistic for goodness-of-fit?
- RQ4Is the computation of the test statistic feasible for large trees, and can it be reduced to a known combinatorial optimization problem?
- RQ5Can the method distinguish between different tree-generating processes even when their mean trees are identical?
Key findings
- The empirical d-mean converges almost surely to the true d-mean under compactness and uniqueness, establishing consistency of the estimator.
- An invariance principle holds for the process (gₙ(y) − g(y)), implying the asymptotic distribution of the supremum deviation is known.
- The test statistic sup_y |gₙ(y) − g(y)| is distribution-free under the null and suitable for Kolmogorov-Smirnov-type testing.
- The computation of the supremum is reduced to a minimal cut problem in a derived network, enabling polynomial-time computation.
- The method successfully distinguishes between different Galton-Watson processes and classifies FGF protein families, even when mean trees are identical.
- The space of trees under the given metric is not of negative curvature, which distinguishes it from CAT(0) spaces like phylogenetic tree spaces.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.