[Paper Review] Estimating Fold Changes from Partially Observed Outcomes with Applications in Microbial Metagenomics
This paper develops a method to estimate fold-changes in the mean abundance of multiple taxa when only partially observed outputs are available, addressing sample- and category-specific perturbations in microbial metagenomics. It provides identifiability via constraints, a penalized estimation approach, robust hypothesis testing, and an application to colorectal cancer meta-analysis.
We consider the problem of estimating fold-changes in the expected value of a multivariate outcome observed with unknown sample-specific and category-specific perturbations. This challenge arises in high-throughput sequencing studies of the abundance of microbial taxa because microbes are systematically over- and under-detected relative to their true abundances. Our model admits a partially identifiable estimand, and we establish full identifiability by imposing interpretable parameter constraints. To reduce bias and guarantee the existence of estimators in the presence of sparse observations, we apply an asymptotically negligible and constraint-invariant penalty to our estimating function. We develop a fast coordinate descent algorithm for estimation, and an augmented Lagrangian algorithm for estimation under null hypotheses. We construct a model-robust score test and demonstrate valid inference even for small sample sizes and violated distributional assumptions. The flexibility of the approach and comparisons to related methods are illustrated through a meta-analysis of microbial associations with colorectal cancer.
Motivation & Objective
- Motivate and formalize the problem of estimating fold-differences in the mean of a nonnegative multivariate outcome observed with sample- and category-specific perturbations.
- Develop identifiability through parameter constraints and an interpretable estimand on fold-differences of true abundances.
- Propose a fast estimation algorithm with Firth-type bias reduction and a constraint-driven identifiability approach.
- Develop model-robust inference procedures, including a robust score test and robust Wald tests, with performance under small samples and distributional misspecification.
- Demonstrate the method via simulations and a meta-analysis of colorectal cancer-associated microbiomes.
Proposed method
- Specify a log-linear model for the unobserved true abundances and a perturbed, partially observed version with unknown sample-specific and taxon-specific effects.
- Establish partial identifiability and define equivalence classes of parameters; impose a smooth identifiability constraint (pseudo-Huber) to identify fold-differences.
- Use a penalized likelihood with a Firth-type penalty to ensure finite estimates under separation and sparsity; solve via coordinate descent and augmented data techniques.
- Derive a model-robust score test for hypotheses on log-fold differences under identifiability constraints; provide a robust score statistic and an alternative robust Wald test.
- Provide an augmented Lagrangian optimization framework for constrained estimation under null hypotheses; ensure invariance of the penalized likelihood to constraint choices.

Experimental results
Research questions
- RQ1How can one estimate fold-differences in the mean of true abundances when observations are distorted by unknown sample- and category-specific effects?
- RQ2What identifiability constraints are needed to meaningfully interpret such fold-changes in microbiome data?
- RQ3Can we develop fast, bias-reducing estimation procedures and robust inference methods that perform well with small samples and possible distributional misspecification?
- RQ4How does the proposed method perform in simulations under Poisson and zero-inflated negative binomial settings and in a real colorectal cancer microbiome meta-analysis?
Key findings
- The method identifies equivalence classes of the log-mean parameter and achieves full identifiability by imposing a constraint on the rows of the coefficient matrix.
- A Firth-penalized profile likelihood with a coordinate descent algorithm provides stable estimation under sparse observations and potential separation.
- Robust score tests control Type I error well in larger samples and remain conservative in small samples, while robust Wald tests can be anti-conservative in small samples.
- Power is higher under Poisson data than ZINB, and increases with sample size and, to a lesser extent, with the number of taxa.
- In a colorectal cancer meta-analysis using Wirbel et al. data, 30 taxa showed differential abundance associated with CRC status under a 0.1 FDR threshold using robust score testing.
- The approach yields interpretable log-fold differences in true abundances across covariate levels, relative to typical differences across taxa.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.