[Paper Review] Bayesian biclustering for microbial metagenomic sequencing data via multinomial matrix factorization
This paper proposes a Bayesian multinomial matrix factorization model with a phylogenetic Indian buffet process prior to jointly cluster microbes and hosts in metagenomic data, accounting for compositionality, sparsity, and over-dispersion. It successfully identifies biologically relevant, overlapping microbial communities linked to inflammatory bowel disease, including known families like Bacteroidaceae and Enterobacteriaceae, with improved interpretability through taxonomic tree integration.
High-throughput sequencing technology provides unprecedented opportunities to quantitatively explore human gut microbiome and its relation to diseases. Microbiome data are compositional, sparse, noisy, and heterogeneous, which pose serious challenges for statistical modeling. We propose an identifiable Bayesian multinomial matrix factorization model to infer overlapping clusters on both microbes and hosts. The proposed method represents the observed over-dispersed zero-inflated count matrix as Dirichlet-multinomial mixtures on which latent cluster structures are built hierarchically. Under the Bayesian framework, the number of clusters is automatically determined and available information from a taxonomic rank tree of microbes is naturally incorporated, which greatly improves the interpretability of our findings. We demonstrate the utility of the proposed approach by comparing to alternative methods in simulations. An application to a human gut microbiome dataset involving patients with inflammatory bowel disease reveals interesting clusters, which contain bacteria families Bacteroidaceae, Bifidobacteriaceae, Enterobacteriaceae, Fusobacteriaceae, Lachnospiraceae, Ruminococcaceae, Pasteurellaceae, and Porphyromonadaceae that are known to be related to the inflammatory bowel disease and its subtypes according to biological literature. Our findings can help generate potential hypotheses for future investigation of the heterogeneity of the human gut microbiome.
Motivation & Objective
- To address the challenges of compositional, sparse, heterogeneous, and noisy microbiome data in statistical modeling.
- To develop a joint clustering framework that simultaneously identifies overlapping microbial and host clusters.
- To incorporate taxonomic hierarchy information to enhance biological interpretability of inferred clusters.
- To enable automatic cluster determination and full posterior inference under a hierarchical Bayesian model.
- To improve detection of disease-relevant microbial communities in inflammatory bowel disease (IBD) patients.
Proposed method
- Models microbiome count data as Dirichlet-multinomial mixtures to account for over-dispersion and zero-inflation.
- Uses a phylogenetic Indian buffet process (pIBP) prior to encode taxonomic relationships into latent cluster structures.
- Employs a hierarchical Bayesian framework to jointly infer microbial and host clusters with overlapping memberships.
- Applies sparse matrix factorization to represent the observed count matrix via latent binary indicators Z.
- Incorporates host-specific parameters s_ij and t_ij to allow for heterogeneous cluster assignments across individuals.
- Uses MCMC for full posterior inference, enabling uncertainty quantification and probabilistic cluster characterization.
Experimental results
Research questions
- RQ1Can a Bayesian multinomial matrix factorization model effectively handle the compositional and zero-inflated nature of metagenomic sequencing data?
- RQ2How does incorporating phylogenetic tree information improve the biological interpretability of biclusters in microbiome data?
- RQ3What is the impact of prior knowledge on cluster discovery accuracy and reproducibility in the absence of a gold standard?
- RQ4Can the model automatically determine the optimal number of overlapping clusters without pre-specification?
- RQ5How does the proposed method compare to alternative approaches in detecting IBD-associated microbial communities?
Key findings
- The method successfully identified four distinct biclusters in the IBD dataset, with one cluster enriched for patients with inflammatory bowel disease.
- The model detected known IBD-associated bacterial families, including Bacteroidaceae, Bifidobacteriaceae, Enterobacteriaceae, Fusobacteriaceae, Lachnospiraceae, Ruminococcaceae, Pasteurellaceae, and Porphyromonadaceae.
- Incorporating taxonomic tree information significantly improved cluster interpretability, as evidenced by a higher log probability (-91.48 vs. -119.95) of the inferred matrix under the tree structure.
- Without tree priors, the method failed to identify the IBD-related cluster and showed poor taxonomic coherence within clusters.
- The posterior mean of the latent assignment matrix Z from the proposed model best captured the underlying abundance patterns compared to alternative priors and deterministic thresholds.
- The method demonstrated superior performance in simulations and real data, particularly in recovering biologically meaningful clusters when prior knowledge was used.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.