Skip to main content
QUICK REVIEW

[论文解读] Metagenomic Analysis using Phylogenetic Placement -- A Review of the First Decade

Lucas Czech, Alexandros Stamatakis|arXiv (Cornell University)|Feb 7, 2022
Genomics and Phylogenetic Studies参考文献 296被引用 59
一句话总结

本综述总结了宏基因组学中系统发育定位技术过去十年的发展,提出了一套从原始序列到可发表结果的完整工作流程。它强调了将查询序列定位到参考系统发育树上,可提升分类学鉴定、多样性估算和生态推断的准确性,同时指出了潜在问题并实现了与环境元数据的整合。

ABSTRACT

Phylogenetic placement refers to a family of tools and methods to analyze, visualize, and interpret the tsunami of metagenomic sequencing data generated by high-throughput sequencing. Compared to alternative (e. g., similarity-based) methods, it puts metabarcoding sequences into a phylogenetic context using a set of known reference sequences and taking evolutionary history into account. Thereby, one can increase the accuracy of metagenomic surveys and eliminate the requirement for having exact or close matches with existing sequence databases. Phylogenetic placement constitutes a valuable analysis tool per se, but also entails a plethora of downstream tools to interpret its results. A common use case is to analyze species communities obtained from metagenomic sequencing, for example via taxonomic assignment, diversity quantification, sample comparison, and identification of correlations with environmental variables. In this review, we provide an overview over the methods developed during the first ten years. In particular, the goals of this review are (i) to motivate the usage of phylogenetic placement and illustrate some of its use cases, (ii) to outline the full workflow, from raw sequences to publishable figures, including best practices, (iii) to introduce the most common tools and methods and their capabilities, (iv) to point out common placement pitfalls and misconceptions,(v) to showcase typical placement-based analyses, and how they can help to analyze, visualize, and interpret phylogenetic placement data.

研究动机与目标

  • 提供对过去十年间在宏基因组学中应用的系统发育定位方法的系统性概述。
  • 通过展示其在分类学鉴定和进化背景解析方面相对于相似性比对方法的优势,推动系统发育定位的采用。
  • 概述从序列处理到结果可视化与解释的端到端工作流程的最佳实践。
  • 识别并澄清系统发育定位分析中的常见误解和方法论陷阱。
  • 展示如何利用定位数据开展下游分析,以实现生态推断,包括多样性估算、样本比较以及环境元数据整合。

提出的方法

  • 使用最大似然法(ML)计算查询序列在参考系统发育树(RT)各分支上的似然值,实现概率性定位。
  • 采用已知参考序列(RSs)的参考比对(RA)来推断RT,并支持定位似然值的计算。
  • 应用似然权重比(LWR)量化定位置信度,每个查询序列在所有分支上的LWR值总和为1。
  • 利用边缘相关性(Edge Correlation)和定位因子分解(Placement-Factorization)等方法,将定位数据与环境元数据整合,检测类群丰度与环境变量之间的关联。
  • 在定位分布上应用聚类与排序技术(如k-means、层次聚类),以可视化群落组成和样本间关系。
  • 在定位因子分解中使用广义线性模型(GLMs)来建模多种元数据类型(数值型、二值型、分类变量)与类群水平丰度模式之间的关系。

实验结果

研究问题

  • RQ1与基于相似性的方法相比,系统发育定位在宏基因组调查中如何提升分类学鉴定的准确性?
  • RQ2从原始测序读段构建完整系统发育定位工作流程的关键步骤和最佳实践是什么?
  • RQ3如何利用定位数据推断多个样本中物种多样性、群落组成和生态模式?
  • RQ4系统发育定位中的主要误差来源和误解是什么,如何加以缓解?
  • RQ5在何种方式下,系统发育定位数据可与环境元数据有意义地关联,以揭示微生物群落结构的生态驱动因素?

主要发现

  • 系统发育定位通过整合进化历史,显著提升了分类学鉴定的准确性,减少了基于相似性方法带来的假阳性结果。
  • 使用高质量参考序列构建的参考树和参考比对,可在查询序列在数据库中缺乏近缘匹配时,仍实现稳健的定位。
  • 定位因子分解成功识别出丰度与环境变量相关的类群,实现了对生态数据中嵌套依赖关系的检测。
  • 边缘相关性有效可视化了群落丰度模式与环境变量(如pH或温度)强相关联的树区域。
  • 该方法框架通过基于元数据驱动信号将树分解为嵌套类群,支持降维和样本排序。
  • 尽管具有优势,该方法仍缺乏标准化的指标来评估定位质量,特别是在区分缺失的参考序列与尚未描述的新类群方面。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。