Skip to main content
QUICK REVIEW

[论文解读] Multi-sample Estimation of Bacterial Composition Matrix in Metagenomics Data

Yuanpei Cao, Anru R. Zhang|arXiv (Cornell University)|Jun 7, 2017
Bayesian Methods and Mixture Models参考文献 45被引用 8
一句话总结

本文提出了一种核范数正则化的最大似然估计器,通过利用多个样本间的低秩结构,改进了宏基因组学中细菌组成矩阵的估计。通过使用泊松-多项分布建模稀疏计数数据,并采用邻近梯度优化方法,该方法减少了因零膨胀的稀有分类群带来的偏差,在Kullback-Leibler散度和Frobenius范数下实现了极小极大最优误差界。

ABSTRACT

Metagenomics sequencing is routinely applied to quantify bacterial abundances in microbiome studies, where the bacterial composition is estimated based on the sequencing read counts. Due to limited sequencing depth and DNA dropouts, many rare bacterial taxa might not be captured in the final sequencing reads, which results in many zero counts. Naive composition estimation using count normalization leads to many zero proportions, which tend to result in inaccurate estimates of bacterial abundance and diversity. This paper takes a multi-sample approach to the estimation of bacterial abundances in order to borrow information across samples and across species. Empirical results from real data sets suggest that the composition matrix over multiple samples is approximately low rank, which motivates a regularized maximum likelihood estimation with a nuclear norm penalty. An efficient optimization algorithm using the generalized accelerated proximal gradient and Euclidean projection onto simplex space is developed. The theoretical upper bounds and the minimax lower bounds of the estimation errors, measured by the Kullback-Leibler divergence and the Frobenius norm, are established. Simulation studies demonstrate that the proposed estimator outperforms the naive estimators. The method is applied to an analysis of a human gut microbiome dataset.

研究动机与目标

  • 解决由于测序深度有限和DNA丢失导致的零膨胀、稀疏的宏基因组计数数据问题。
  • 克服朴素组成估计方法带来的不准确性,避免由此引发的多样性估计偏差和下游分析问题。
  • 利用多样本信息,通过挖掘真实组成矩阵的近似低秩结构来改进估计。
  • 在泊松-多项分布模型下,构建具有理论保证的正则化最大似然框架,以控制估计误差。
  • 建立估计误差的极小极大下界和上界,证明所提估计器的最优性。

提出的方法

  • 采用泊松-多项分布模型建模问题:每个样本的总读长服从泊松分布,各类群的计数服从多项分布。
  • 提出一种核范数正则化的最大似然估计器,以在组成矩阵上强制实现低秩结构。
  • 采用广义加速邻近梯度算法,并结合单纯形上的欧氏投影求解优化问题。
  • 引入约束集 D(T) 以在Kullback-Leibler散度和Frobenius范数下控制估计误差。
  • 通过对称化和压缩论证推导估计误差的集中不等式。
  • 基于经验过程理论推导理论界,包括估计误差的上界和极小极大下界。

实验结果

研究问题

  • RQ1能否有效利用多样本信息,以改善稀疏宏基因组数据中细菌组成的估计?
  • RQ2在多个样本中,真实组成矩阵是否表现出低秩结构,从而支持核范数正则化?
  • RQ3与朴素归一化方法相比,所提正则化估计器在估计精度和多样性估计方面表现如何?
  • RQ4在Kullback-Leibler散度和Frobenius范数下,所提估计器的理论误差界(上界与极小极大下界)是什么?
  • RQ5在泊松-多项分布模型下,所提方法是否为极小极大最优?

主要发现

  • 在泊松-多项分布模型下,所提估计器在Kullback-Leibler散度和Frobenius范数下均达到了极小极大最优收敛速率。
  • 理论分析表明,估计误差以高概率有界,且上界与极小极大下界仅相差对数因子。
  • 模拟研究显示,所提方法在稀有分类群和整体组成估计方面显著优于朴素归一化方法。
  • 在人类肠道微生物组数据集上的实证分析验证了组成矩阵的低秩结构,并证实了该方法的实际有效性。
  • 该方法有效降低了因采样不足和DNA丢失导致的零计数带来的偏差,改善了下游多样性估计。
  • 理论界在最小假设下推导得出,明确依赖于样本量 N、分类群数量 p 和真实组成矩阵的秩 r。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。