Skip to main content
QUICK REVIEW

[论文解读] Progressive Mauve: Multiple alignment of genomes with gene flux and rearrangement

Aaron E. Darling, Bob Mau|ArXiv.org|Oct 30, 2009
Genomics and Phylogenetic Studies参考文献 58被引用 6
一句话总结

Progressive Mauve 提出了一种新颖的多基因组比对方法,通过使用成对断点得分和概率过滤,能够准确比对经历广泛基因漂移、重排及插入缺失事件的基因组。该方法在比对23株肠杆菌科基因组时,准确度优于以往方法,揭示了与调控分化相关的广泛非编码区变异,并定义了核心基因组2.46 Mbp和泛基因组15.2 Mbp。

ABSTRACT

Multiple genome alignment remains a challenging problem. Effects of recombination including rearrangement, segmental duplication, gain, and loss can create a mosaic pattern of homology even among closely related organisms. We describe a method to align two or more genomes that have undergone large-scale recombination, particularly genomes that have undergone substantial amounts of gene gain and loss (gene flux). The method utilizes a novel alignment objective score, referred to as a sum-of-pairs breakpoint score. We also apply a probabilistic alignment filtering method to remove erroneous alignments of unrelated sequences, which are commonly observed in other genome alignment methods. We describe new metrics for quantifying genome alignment accuracy which measure the quality of rearrangement breakpoint predictions and indel predictions. The progressive genome alignment algorithm demonstrates markedly improved accuracy over previous approaches in situations where genomes have undergone realistic amounts of genome rearrangement, gene gain, loss, and duplication. We apply the progressive genome alignment algorithm to a set of 23 completely sequenced genomes from the genera Escherichia, Shigella, and Salmonella. The 23 enterobacteria have an estimated 2.46Mbp of genomic content conserved among all taxa and total unique content of 15.2Mbp. We document substantial population-level variability among these organisms driven by homologous recombination, gene gain, and gene loss. Free, open-source software implementing the described genome alignment approach is available from http://gel.ahabs.wisc.edu/mauve .

研究动机与目标

  • 解决在广泛存在基因获得、丢失、重复和重排的情况下,实现准确多基因组比对的挑战。
  • 开发一种方法,以区分经历重组和基因漂移的基因组中的同源基因、异源同源基因和非同源序列。
  • 通过新型评分函数整合共线性图谱与比对,提升比对准确度。
  • 利用断点和插入缺失预测准确度的度量指标,量化比对质量。
  • 通过生成跨多样化微生物菌株的可靠全基因组比对,支持比较基因组学和群体基因组学研究。

提出的方法

  • 该方法采用新型成对断点得分,惩罚引入重排或锚定于重复区域的比对配置。
  • 采用渐进式比对策略,对不同分类群子集的保守区域进行比对,避免依赖单一参考基因组。
  • 通过概率过滤步骤去除非同源序列的错误比对,提升特异性。
  • 在单一框架内整合共线性图谱与比对,实现锚点定位与比对结果的相互优化。
  • 采用基于动态规划的带状评分比对方法,在保持敏感性的同时提升计算效率。
  • 支持通过Java实现的交互式可视化,便于探索基因组中保守与可变区域。

实验结果

研究问题

  • RQ1当基因组经历显著基因漂移和重排时,如何实现准确的多基因组比对?
  • RQ2在密切相关的细菌物种中,由于重组、基因获得和丢失,非编码区在多大程度上表现出可变性?
  • RQ3在复杂进化事件背景下,统一的比对框架能否有效区分同源基因、异源同源基因和非同源序列?
  • RQ4肠杆菌科中核心基因组与泛基因组的多样性程度如何,包括非编码调控区域?
  • RQ5重组在多大程度上影响非编码区的保守性模式,这对调控分化意味着什么?

主要发现

  • 在具有真实基因漂移和重排的模拟数据集上,Progressive Mauve 算法的比对准确度显著优于以往方法。
  • 该方法成功比对了来自大肠杆菌属、志贺氏菌属和沙门氏菌属的23个完整测序基因组,揭示了2.46 Mbp的核心基因组和15.2 Mbp的泛基因组。
  • 观察到显著的群体水平变异,其中大部分变异集中于非编码区,提示存在调控分化。
  • yhjE 基因座区域表现出非编码序列的高变异性,原因包括插入、缺失和重排,某些菌株中甚至出现RIP元件的非典型替换。
  • 非编码区变异的系统发育模式(如RIP元件区域)表明,其分化由重组驱动而非垂直遗传,且在大肠杆菌与志贺氏菌菌株的一个类群中存在趋同进化的证据。
  • 对23重比对的筛查识别出102个严格位于非编码区的区域,其保守性模式高度可变,凸显了非编码区广泛的调控潜力。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。