[论文解读] Elementary methods provide more replicable results in microbial differential abundance analysis
本研究利用来自16S rRNA和宏基因组测序研究的53个数据集,评估了14种差异丰度分析方法在微生境研究中的表现。研究发现,基础方法——如Mann-Whitney U检验、有序回归、相对丰度上的线性回归/t检验以及存在/缺失数据上的逻辑回归——在数据集划分和独立研究之间表现出更高的可重复性,尽管缺乏正式的真值参考,其结果一致性仍优于复杂方法。
Differential abundance analysis is a key component of microbiome studies. Although dozens of methods exist there is currently no consensus on the preferred methods. While the correctness of results in differential abundance analysis is an ambiguous concept and cannot be fully evaluated without setting the ground truth and employing simulated data, we argue that a well-performing method should be effective in producing highly reproducible results. We compared the performance of 14 differential abundance analysis methods by employing datasets from 53 taxonomic profiling studies based on 16S rRNA gene or shotgun metagenomic sequencing. For each method, we examined how the results replicated between random partitions of each dataset and between datasets from separate studies. While certain methods showed good consistency, some widely used methods were observed to produce a substantial number of conflicting findings. Overall, when considering consistency together with sensitivity, the best performance was attained by analyzing relative abundances with a non-parametric method (Wilcoxon test or ordinal regression model) or linear regression/t-test. Moreover, a comparable performance was obtained by analyzing presence/absence of taxa with logistic regression.
研究动机与目标
- 评估差异丰度分析方法在多样化微生物组数据集中的可重复性。
- 识别在相同数据集的随机划分和独立研究中均能产生一致结果的方法。
- 挑战在微生物组研究中复杂、专用方法天然优于基础统计方法的假设。
- 为选择能最大化结果一致性的方法提供基于证据的指导。
提出的方法
- 作者在来自16S rRNA和宏基因组测序的53个分类谱系数据集上评估了14种差异丰度方法。
- 对于每个数据集,通过测量多个随机划分间的一致性来评估其内部可重复性。
- 通过比较来自不同研究的数据集间的结果,评估跨研究的一致性,以检测差异丰度分类群的重叠程度作为指标。
- 相对丰度使用非参数检验(Mann-Whitney U检验)、有序回归、线性回归和t检验进行分析。
- 存在/缺失数据使用逻辑回归分析,以评估其可重复性。
- 性能基于在划分和研究间的一致性进行评估,同时考虑敏感性。
实验结果
研究问题
- RQ1在相同数据集的随机划分中,哪些差异丰度方法能产生最可重复的结果?
- RQ2广泛使用的复杂方法与基础统计方法在跨研究一致性方面如何比较?
- RQ3方法选择在多大程度上影响微生物差异丰度发现的可重复性?
- RQ4在微生物组研究中,简单统计方法是否能在结果一致性方面超越复杂模型?
主要发现
- 基础方法如Mann-Whitney U检验和相对丰度上的有序回归在相同数据集的随机划分中表现出高度一致性。
- 相对丰度上的线性回归和t检验也表现出强可重复性,其表现与更复杂的模型相当。
- 存在/缺失数据上的逻辑回归达到了相似的一致性水平,表明其是二值数据的稳健替代方案。
- 一些广泛使用的方法,包括部分专用微生物组工具,在划分和研究之间产生了大量矛盾结果。
- 综合考虑一致性和敏感性,基础方法整体表现最佳。
- 本研究未发现复杂、基于模型的方法在可重复性方面一致优于基础统计方法的证据。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。