[论文解读] A tree-based model for addressing sparsity and taxa covariance in microbiome compositional count data
本文提出了一种名为逻辑树正态(LTN)的新型贝叶斯生成模型,用于处理微生物组组成数据。该模型通过基于树的二项分解与Pólya-Gamma扩展,结合了对数比正态(LN)模型的灵活协方差结构与基于树的狄利克雷(DT)模型的计算效率。该方法在稀疏性和低秩假设下实现了可扩展的高维推断,在纵向1型糖尿病(T1D)队列数据集上表现出色的关联检验与协方差估计性能。
Microbiome compositional data are often high-dimensional, sparse, and exhibit pervasive cross-sample heterogeneity. Generative modeling is a popular approach to analyze such data, and effective generative models must accurately characterize these key features. While high-dimensionality and abundance of zeros have received much attention, existing models often lack flexibility in capturing complex cross-sample variability. This limitation can affect statistical efficiency and lead to misleading conclusions in tasks like differential abundance analysis, clustering, and network analysis. We introduce a generative model, the "logistic-tree normal" (LTN) model, which addresses this issue and effectively captures key characteristics of microbiome data, including abundance of zeros. LTN employs a tree-based decomposition to aggregate sparse taxa counts and uses a (multivariate) logistic-normal distribution at tree splits, allowing for flexible covariance adjustments among taxa as needed. The latent Gaussian structure of LTN enables the incorporation of multivariate analysis tools that enforce sparsity or low-rank covariance assumptions. As a versatile, fully generative model, LTN supports a wide range of applications and offers efficient Bayesian inference computational recipes through conjugate blocked Gibbs sampling with Pólya-Gamma augmentation. We demonstrate application of LTN in a compositional mixed-effects model for differential abundance analysis using numerical experiments and a reanalysis of the infant cohort in the DIABIMMUNE study. Our findings illustrate that LTN, by adequately accounting for cross-sample heterogeneity, appropriately generates the proportion of zeros without requiring an explicit zero-inflation component, confirming a recent viewpoint that "zero-inflation" in count-based sequencing data are often results of unaccounted cross-sample variation.
研究动机与目标
- 解决现有模型在处理高维微生物组组成数据中复杂协方差与计算可扩展性方面存在的局限性。
- 开发一种生成模型,保留对数比正态(LN)模型丰富的协方差结构,同时通过基于树的分解实现计算可行性。
- 在稀疏性与低秩假设下,实现微生物组关联研究与协方差估计的有效贝叶斯推断。
- 在纵向微生物组数据分析中展示LTN模型的实用性,特别是在检测与疾病风险关联方面。
提出的方法
- LTN模型将多项分布似然在系统发育树的内部节点处分解为一系列二项概率。
- 使用多变量正态分布建模这些二项概率的对数优势比,从而在分类群之间实现灵活的协方差结构。
- 引入Pólya-Gamma辅助变量,通过层次模型中的共轭性实现高效的Gibbs采样。
- 该模型支持对数优势比协方差矩阵的稀疏性与低秩假设,促进高维推断。
- 构建一个通用的混合效应模型,利用LTN对组成随机效应进行建模,并检验与协变量的关联。
- 该框架允许整合先验信息,例如在精度矩阵上应用图稀疏化(graphical Lasso),以推断分类群之间的稀疏相互作用网络。
实验结果
研究问题
- RQ1能否开发一种模型,将对数比正态模型的灵活协方差与基于树模型的计算效率相结合,用于微生物组数据?
- RQ2如何有效将稀疏性与低秩结构融入组成数据模型,以提升可扩展性与可解释性?
- RQ3LTN模型是否能在纵向研究中检测到微生物组组成与疾病风险之间的有意义关联?
- RQ4与现有方法相比,LTN模型在估计微生物分类群之间潜在协方差结构方面的表现如何?
主要发现
- LTN模型成功捕捉了微生物分类群之间的复杂协方差结构,在灵活性上优于传统的狄利克雷-多项分布模型。
- Pólya-Gamma扩展的使用实现了高效的Gibbs采样,使得即使在超过50个分类群的情况下,贝叶斯推断也具备可扩展性。
- 在DIABIMMUNE T1D队列研究中,该模型检测到微生物组组成与母乳喂养、固体食物及大豆制品等饮食因素之间的显著关联。
- PMAPs分析显示,大麦、黑麦及固体食物的引入与已知的菌群失调指标——厚壁菌门/拟杆菌门比值变化相关。
- 模型识别出一条具有稳定相对丰度变化的节点链,与大豆制品摄入相关,提示其对特定分类群(如OTU 4439360)存在累积影响。
- 分析结果确认了已知的生物学模式,例如母乳喂养期间双歧杆菌属的富集,特别是长双歧杆菌与两歧双歧杆菌等物种。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。