Skip to main content
QUICK REVIEW

[论文解读] Bayesian Modeling of Microbiome Data for Differential Abundance Analysis

Qiwei Li, Shuang Jiang|arXiv (Cornell University)|Feb 23, 2019
Gut microbiota and health参考文献 63被引用 7
一句话总结

本文提出了一种贝叶斯分层模型 ZINB-DPP,用于宏基因组数据的差异丰度分析,将零膨胀负二项分布(ZINB)建模与狄利克雷过程先验相结合,实现特征选择和系统发育结构建模。该方法在检测具有过多零值、过度离散和测序深度不均的生物相关类群方面优于现有方法,同时控制了错误发现率并整合了进化关系。

ABSTRACT

The advances of next-generation sequencing technology have accelerated study of the microbiome and stimulated the high throughput profiling of metagenomes. The large volume of sequenced data has encouraged the rise of various studies for detecting differentially abundant taxonomic features across healthy and diseased populations, with the ultimate goal of deciphering the relationship between the microbiome diversity and health conditions. As the microbiome data are high-dimensional, typically featuring by uneven sampling depth, overdispersion and a huge amount of zeros, these data characteristics often hamper the downstream analysis. Moreover, the taxonomic features are implicitly imposed by the phylogenetic tree structure and often ignored. To overcome these challenges, we propose a Bayesian hierarchical modeling framework for the analysis of microbiome count data for differential abundance analysis. Under this framework, we introduce a bi-level Bayesian hierarchical model that allows a flexible choice of the count generating process, and hyperpriors in the feature selection scheme. We particularly focus on employing a zero-inflated negative binomial model with a Bayesian nonparametric prior model on the bottom level, and applying Gaussian mixture models for differentially abundant taxa detection on the top level. Our method allows for the simultaneous modeling of sample heterogeneity and detecting differentially abundant taxa. We conducted comprehensive simulations and summarized the improved statistical performances of the proposed model. We applied the model in two real microbiome study datasets and successfully identified biologically validated differentially abundant taxa. We hope that the proposed framework and model can facilitate further microbiome studies and elucidate disease etiology.

研究动机与目标

  • 解决高维宏基因组计数数据中零值过多、过度离散和测序深度不均的挑战。
  • 开发一种统一的统计框架,通过信息性先验实现基于模型的标准化,避免临时的预标准化处理。
  • 提高在不同疾病状态(如结直肠癌、精神分裂症)下识别差异丰度类群的检测效能和准确性。
  • 利用马氏随机场先验将系统发育关系整合到差异丰度检验中。
  • 通过贝叶斯错误发现率估计控制错误发现率,提升结果的可重复性和生物学相关性。

提出的方法

  • 该模型采用两级分层结构:底层使用 ZINB 分布对具有过多零值和过度离散的宏基因组计数数据进行建模。
  • 通过先验分布中的随机约束实现基于模型的标准化,消除了对预标准化的依赖。
  • 顶层通过具有狄利克雷过程先验的正态分布混合模型实现非参数化特征选择,并自动确定差异丰度类群。
  • 利用马氏随机场先验整合系统发育结构,促使树中相邻类群共享差异丰度状态。
  • 所有参数通过马氏链蒙特卡洛(MCMC)采样进行估计,实现完整的后验推断。
  • 应用贝叶斯错误发现率(FDR)控制,识别显著类群,同时考虑后验包含概率的不确定性。

实验结果

研究问题

  • RQ1贝叶斯分层模型能否在无需预标准化的情况下,有效处理具有零值过多、过度离散和高维特性的宏基因组计数数据?
  • RQ2通过马氏随机场先验引入系统发育结构,如何提升对具有生物学意义的差异丰度类群的检测能力?
  • RQ3与传统方法(如 Kruskal–Wallis、DESeq2、edgeR、metagenomeSeq)相比,ZINB-DPP 模型在统计效能和 FDR 控制方面表现如何?
  • RQ4该模型在真实世界数据集(如结直肠癌和精神分裂症)中,能在多大程度上恢复已知的宏基因组-疾病关联?
  • RQ5该模型能否检测到如具核梭杆菌(Fusobacterium nucleatum)和弯曲杆菌(Campylobacter)等生物学上相关但可能被标准方法遗漏的共现类群?

主要发现

  • 在结直肠癌数据集中,ZINB-DPP 模型在 1% 贝叶斯 FDR 阈值下检测到 10 种差异丰度物种,其中 7 种得到先前生物学证据的支持。
  • ZINB-DPP 模型成功识别出 Synergistaceae 至 Synergistetes 系统发育分支在 CRC 中富集,该发现与先前研究一致。
  • 与 Kruskal–Wallis 方法报告的 12 种物种相比(仅 7 种有生物学支持),ZINB-DPP 模型在 11 种物种中确认了 6 种,表现出更高的精确度。
  • 在精神分裂症研究中,ZINB-DPP 在 5% 贝叶斯 FDR 下识别出 8 种差异丰度类群,其中 5 种与 DESeq2 和 metagenomeSeq 结果重叠,且唯一检测到 Veillonella parvula。
  • DESeq2 和 edgeR 尽管检测数量更多,但 FDR 控制效果较差,FDR 控制方面 ZINB-DPP 表现更优,表现为假阳性更少。
  • ZINB-DPP 模型检测到了具核梭杆菌与弯曲杆菌之间的共现模式,而 metagenomeSeq 和 DM 模型未能检测到,凸显其对生物共现关系的更高敏感性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。