[Paper Review] Non-bifurcating phylogenetic tree inference via the adaptive LASSO
This paper proposes an adaptive LASSO-based regularization method for non-bifurcating phylogenetic tree inference, enabling detection of zero-length branches indicative of polytomies and sampled ancestors. By applying $β$-adaptive LASSO penalties to branch lengths, the method achieves topological consistency and outperforms thresholding and non-adaptive LASSO in sparsity recovery and computational efficiency.
Phylogenetic tree inference using deep DNA sequencing is reshaping our understanding of rapidly evolving systems, such as the within-host battle between viruses and the immune system. Densely sampled phylogenetic trees can contain special features, including <i>sampled ancestors</i> in which we sequence a genotype along with its direct descendants, and <i>polytomies</i> in which multiple descendants arise simultaneously. These features are apparent after identifying zero-length branches in the tree. However, current maximum-likelihood based approaches are not capable of revealing such zero-length branches. In this article, we find these zero-length branches by introducing adaptive-LASSO-type regularization estimators for the branch lengths of phylogenetic trees, deriving their properties, and showing regularization to be a practically useful approach for phylogenetics. Supplementary materials for this article are available online.
Motivation & Objective
- Address the limitation of current maximum-likelihood methods in detecting zero-length branches, such as polytomies and sampled ancestors, in phylogenetic trees.
- Develop a penalized likelihood framework that encourages sparsity in branch lengths to identify non-bifurcating topologies.
- Establish theoretical consistency of the adaptive phylogenetic LASSO under mild regularity conditions.
- Provide a computationally efficient optimization algorithm based on proximal gradient methods for solving the non-convex, non-smooth regularization problem.
- Demonstrate superior performance compared to heuristic thresholding and non-adaptive LASSO in synthetic and real data experiments.
Proposed method
- Formulate a penalized maximum-likelihood estimation problem with an adaptive LASSO penalty on branch lengths to induce sparsity.
- Use the adaptive LASSO penalty with weights inversely proportional to initial estimates of branch length coefficients to improve variable selection.
- Apply proximal gradient descent with FISTA-style acceleration to solve the non-smooth, non-convex optimization problem arising from the $β$-adaptive LASSO formulation.
- Derive theoretical consistency results for the adaptive phylogenetic LASSO under mild regularity conditions, including topological consistency.
- Integrate the method with standard maximum-likelihood phylogenetic inference pipelines to allow discovery of non-bifurcating topologies.
- Implement a two-step procedure: first estimate branch lengths via penalized likelihood, then use the zero-estimated branches to infer multifurcating or sampled-ancestor topologies.
Experimental results
Research questions
- RQ1Can adaptive LASSO regularization detect zero-length branches in phylogenetic trees, such as polytomies and sampled ancestors, more effectively than existing methods?
- RQ2Does the adaptive phylogenetic LASSO achieve topological consistency in recovering non-bifurcating tree structures under mild regularity conditions?
- RQ3How does the performance of the adaptive LASSO compare to heuristic thresholding and non-adaptive LASSO in sparsity recovery and branch detection accuracy?
- RQ4Can the proposed method achieve computational efficiency comparable to maximum-likelihood inference while enabling discovery of non-bifurcating topologies?
- RQ5What is the impact of adaptive weighting in the LASSO penalty on the identification of true zero-length branches in complex tree topologies?
Key findings
- The adaptive phylogenetic LASSO achieves topological consistency, enabling reliable detection of non-bifurcating tree topologies with polytomies and sampled ancestors.
- In synthetic experiments, the adaptive LASSO significantly outperforms the non-adaptive LASSO in sparsity recovery, correctly identifying zero-length branches with higher accuracy.
- The method detects short branches with higher sensitivity than thresholding-based approaches across a range of threshold values.
- Compared to rjMCMC, the adaptive LASSO identifies zero-length branches with higher precision and is computationally more efficient.
- The adaptive LASSO maintains consistent performance across varying sequence lengths and mutation rates, demonstrating robustness in diverse simulation settings.
- The method is effective even when zero-length branches have high likelihoods, outperforming non-negative constrained likelihood maximization in challenging topological configurations.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.