Skip to main content
QUICK REVIEW

[论文解读] Bayesian biclustering for microbial metagenomic sequencing data via multinomial matrix factorization

Fangting Zhou, Kejun He|arXiv (Cornell University)|May 17, 2020
Gut microbiota and health参考文献 20被引用 4
一句话总结

该论文提出了一种具有系统发育印度餐厅餐牌过程先验的贝叶斯多项式矩阵分解模型,用于在宏基因组数据中联合聚类微生物和宿主,同时考虑了组成性、稀疏性和过度离散性。该方法成功识别出与炎症性肠病相关的生物上合理的重叠微生物群落,包括已知的菌群如拟杆菌科和肠杆菌科,并通过整合分类树提高了可解释性。

ABSTRACT

High-throughput sequencing technology provides unprecedented opportunities to quantitatively explore human gut microbiome and its relation to diseases. Microbiome data are compositional, sparse, noisy, and heterogeneous, which pose serious challenges for statistical modeling. We propose an identifiable Bayesian multinomial matrix factorization model to infer overlapping clusters on both microbes and hosts. The proposed method represents the observed over-dispersed zero-inflated count matrix as Dirichlet-multinomial mixtures on which latent cluster structures are built hierarchically. Under the Bayesian framework, the number of clusters is automatically determined and available information from a taxonomic rank tree of microbes is naturally incorporated, which greatly improves the interpretability of our findings. We demonstrate the utility of the proposed approach by comparing to alternative methods in simulations. An application to a human gut microbiome dataset involving patients with inflammatory bowel disease reveals interesting clusters, which contain bacteria families Bacteroidaceae, Bifidobacteriaceae, Enterobacteriaceae, Fusobacteriaceae, Lachnospiraceae, Ruminococcaceae, Pasteurellaceae, and Porphyromonadaceae that are known to be related to the inflammatory bowel disease and its subtypes according to biological literature. Our findings can help generate potential hypotheses for future investigation of the heterogeneity of the human gut microbiome.

研究动机与目标

  • 解决在统计建模中宏基因组数据存在的组成性、稀疏性、异质性和噪声性等挑战。
  • 开发一种联合聚类框架,同时识别重叠的微生物和宿主聚类。
  • 整合分类分类信息以增强推断聚类的生物可解释性。
  • 在分层贝叶斯模型下实现自动聚类确定与完整的后验推断。
  • 提高在炎症性肠病(IBD)患者中检测与疾病相关微生物群落的能力。

提出的方法

  • 将微生物组计数数据建模为狄利克雷-多项式混合分布,以处理过度离散性和零膨胀问题。
  • 使用系统发育印度餐厅餐牌过程(pIBP)先验,将分类关系编码到潜在聚类结构中。
  • 采用分层贝叶斯框架,联合推断具有重叠成员关系的微生物和宿主聚类。
  • 应用稀疏矩阵分解,通过潜在二值指示变量Z表示观测到的计数矩阵。
  • 引入个体特异性参数s_ij和t_ij,以允许在不同个体间存在异质的聚类分配。
  • 使用MCMC进行完整后验推断,实现不确定性量化与概率性聚类表征。

实验结果

研究问题

  • RQ1贝叶斯多项式矩阵分解模型能否有效处理宏基因组测序数据的组成性和零膨胀特性?
  • RQ2整合系统发育树信息在多大程度上提升了微生物组数据中双聚类的生物可解释性?
  • RQ3在缺乏金标准的情况下,先验知识对聚类发现的准确性和可重复性有何影响?
  • RQ4该模型能否在无需预先指定的情况下自动确定最优的重叠聚类数量?
  • RQ5与替代方法相比,该方法在检测与IBD相关的微生物群落方面表现如何?

主要发现

  • 该方法在IBD数据集中成功识别出四个不同的双聚类,其中一个聚类在炎症性肠病患者中显著富集。
  • 该模型检测到已知的与IBD相关的细菌家族,包括拟杆菌科、双歧杆菌科、肠杆菌科、梭杆菌科、毛螺菌科、瘤胃球菌科、副嗜血杆菌科和卟啉单胞菌科。
  • 整合分类树信息显著提升了聚类的可解释性,表现为在树结构下推断矩阵的对数概率更高(-91.48 vs. -119.95)。
  • 在无树先验的情况下,该方法未能识别出与IBD相关的聚类,且聚类内部的分类一致性较差。
  • 与替代先验和确定性阈值相比,所提出模型的潜在分配矩阵Z的后验均值最能捕捉到潜在的丰度模式。
  • 在模拟数据和真实数据中,该方法均表现出优越性能,尤其在利用先验知识时,能更有效地恢复具有生物意义的聚类。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。