[论文解读] Testing for differential abundance in compositional counts data, with application to microbiome studies
本文提出DACOMP,一种非参数方法,用于检测具有零膨胀计数和组成性偏差的组成性微生物组数据中的差异丰度。通过利用假设为非差异丰度的参考分类群,DACOMP即使在零率较高时也能提供有效的推断,相较于现有方法在模拟和真实数据集(包括克罗恩病、人类微生物组计划和加标实验)中显著降低了假阳性率。
Identifying which taxa in our microbiota are associated with traits of interest is important for advancing science and health. However, the identification is challenging because the measured vector of taxa counts (by amplicon sequencing) is compositional, so a change in the abundance of one taxon in the microbiota induces a change in the number of sequenced counts across all taxa. The data is typically sparse, with zero counts present either due to biological variance or limited sequencing depth (technical zeros). For low abundance taxa, the chance for technical zeros is non-negligible. We show that existing methods designed to identify differential abundance for compositional data may have an inflated number of false positives due to improper handling of the zero counts. We introduce a novel non-parametric approach which provides valid inference even when the fraction of zero counts is substantial. Our approach uses a set of reference taxa that are non-differentially abundant, which can be estimated from the data or from outside information. We show the usefulness of our approach via simulations, as well as on three different data sets: a Crohn's disease study, the Human Microbiome Project, and an experiment with 'spiked-in' bacteria.
研究动机与目标
- 解决由于在组成性微生物组数据中对零计数处理不当而导致的假阳性率升高的挑战。
- 开发一种在零计数比例较高时(尤其是低丰度分类群)仍能保持有效统计推断的方法。
- 提供一种稳健的非参数方法,利用参考分类群校正组成性偏差,而无需假设参数分布。
- 在具有已知差异丰度模式的真实世界微生物组数据集(包括加标实验)上评估该方法的性能。
- 提供一个公开可用的R包(dacomp),以促进在微生物组研究中的实际应用。
提出的方法
- 该方法使用一组假设为非差异丰度的参考分类群,可从数据中估计或外部提供。
- 对每个待测分类群与参考分类群的相对丰度应用非参数检验(例如,Mann-Whitney U检验或Spearman等级相关)。
- 检验统计量基于每个分类群与参考分类群的对数比值和感兴趣分组或连续性状之间的等级相关性。
- 通过使用参考分类群定义组成性基线实现归一化,从而减少由组成性引起的虚假相关性。
- 该方法通过将零计数视为结构性或技术性零来处理,无需进行插补或变换,避免数据失真。
- 使用Benjamini-Hochberg程序在错误发现率q=0.1的水平上进行多重检验校正。
实验结果
研究问题
- RQ1非参数方法能否在具有高零膨胀和组成性偏差的微生物组数据中有效检测差异丰度?
- RQ2与忽略组成性的方法相比,使用参考分类群是否能提高差异丰度检测的有效性?
- RQ3与ALDEx2、TSS、CSS和CLR变换等现有方法相比,DACOMP在假阳性控制方面表现如何?
- RQ4在已知真实差异丰度的加标实验中,DACOMP能否可靠地检测出已知的差异丰度分类群?
- RQ5在克罗恩病研究和人类微生物组计划等真实数据集中,该方法是否能在控制I类错误的同时保持统计统计功效?
主要发现
- 与ALDEx2、TSS、CSS和CLR变换等现有方法相比,DACOMP显著降低了假阳性率,尤其在零计数占比较高的情况下。
- 在加标实验中,DACOMP以高精度正确识别出两种差异丰度分类群(Rhizobium radiobacter和Alicyclobacillus acidiphilus),而其他方法则表现出较高的假阳性率。
- 即使零计数比例高达90%,该方法仍能保持有效的I类错误控制,优于参数方法和基于变换的方法。
- 在克罗恩病数据集中,DACOMP检测到的生物相关分类群特异性高于ALDEx2和其他方法。
- 人类微生物组计划数据集的分析结果表明,与替代方法相比,DACOMP更一致地识别出与宿主表型相关的微生物关联。
- R包 'dacomp' 已成功实现并公开发布,使得该方法在微生物组研究中可复现且易于应用。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。