Skip to main content
QUICK REVIEW

[Paper Review] Robust Differential Abundance Test in Compositional Data

Shulei Wang|arXiv (Cornell University)|Jan 21, 2021
Geochemistry and Geologic Mapping4 citations
TL;DR

This paper proposes the Robust Differential Abundance (RDB) test, a novel method for differential abundance analysis in compositional data that is robust to zero counts and accounts for the constant-sum constraint. By formulating testing as a reference-based hypothesis and iteratively applying Welch's t-tests on renormalized proportions, the RDB test controls the family-wise error rate asymptotically and enables reliable detection of differentially abundant components without pseudo-count imputation.

ABSTRACT

Differential abundance tests in compositional data are essential and fundamental tasks in various biomedical applications, such as single-cell, bulk RNA-seq, and microbiome data analysis. However, because of the compositional constraint and the prevalence of zero counts in the data, differential abundance analysis in compositional data remains a complicated and unsolved statistical problem. This study introduces a new differential abundance test, the robust differential abundance (RDB) test, to address these challenges. Compared with existing methods, the RDB test is simple and computationally efficient, is robust to prevalent zero counts in compositional datasets, can take the data's compositional nature into account, and has a theoretical guarantee of controlling false discoveries in a general setting. Furthermore, in the presence of observed covariates, the RDB test can work with the covariate balancing techniques to remove the potential confounding effects and draw reliable conclusions. Finally, we apply the new test to several numerical examples using simulated and real datasets to demonstrate its practical merits.

Motivation & Objective

  • To address the persistent challenge of differential abundance testing in compositional data with prevalent zero counts and compositional constraints.
  • To develop a method that is computationally efficient, robust to zero-inflated data, and accounts for the constant-sum nature of compositional data.
  • To provide theoretical guarantees for family-wise error rate control in a general setting.
  • To enable reliable inference in the presence of confounding covariates through integration with covariate balancing techniques.
  • To eliminate the need for ad hoc zero-count imputation strategies like pseudo-counts, which can inflate false discovery rates.

Proposed method

  • Formulates differential abundance testing as a reference-based hypothesis test, where components are compared relative to an unknown reference set.
  • Uses iterative two-stage procedure: first, aggregate t-statistics to determine testing direction; second, apply component-wise directional two-sample t-tests.
  • Applies Welch’s t-test on renormalized proportions in each iteration, avoiding direct handling of zero counts.
  • Employs empirical Bayesian-inspired strategy by iteratively refining test statistics using directional comparisons.
  • Theoretical foundation relies on concentration inequalities and bracketing entropy to establish asymptotic family-wise error rate control.
  • Integrates with covariate balancing techniques to adjust for confounding effects in observational studies.

Experimental results

Research questions

  • RQ1Can a differential abundance test be developed that is robust to zero counts without requiring pseudo-count imputation?
  • RQ2Can such a test maintain theoretical control over the family-wise error rate in a general compositional data setting?
  • RQ3Can the method effectively identify differentially abundant components under the constant-sum constraint?
  • RQ4How does the RDB test perform in the presence of confounding covariates in observational data?
  • RQ5Is the iterative two-stage strategy using renormalized t-statistics both computationally efficient and statistically valid?

Key findings

  • The RDB test asymptotically controls the family-wise error rate at level α under general conditions.
  • The method achieves complete recovery of differentially abundant components with high probability when the signal-to-noise ratio is sufficiently large.
  • Theoretical analysis establishes that the method's error control holds despite dependencies induced by renormalization and iterative procedures.
  • The RDB test avoids the need for pseudo-counts, thereby reducing the risk of false-positive discoveries associated with such imputation.
  • Empirical evaluation on simulated and real datasets demonstrates the method’s practical advantages in power and robustness.
  • Integration with covariate balancing techniques enables reliable inference in the presence of confounding variables.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.