[论文解读] Synteny in Bacterial Genomes: Inference, Organization and Evolution
本研究提出一种基于简约性的方法,用于推断1,108个细菌基因组中的共线性单元(syntons),揭示出小syntons表现出普遍的指数大小分布。作者证明,选择性基因聚集而非随机解体驱动了这些syntons的形成,为塑造细菌基因组组织的共同进化机制提供了证据。
Genes are not located randomly along genomes. Synteny, the conservation of their relative positions in genomes of different species, reflects fundamental constraints on natural evolution. We present approaches to infer pairs of co-localized genes from multiple genomes, describe their organization, and study their evolutionary history. In bacterial genomes, we thus identify synteny units, or "syntons", which are clusters of proximal genes that encompass and extend operons. The size distribution of these syntons divide them into large syntons, which correspond to fundamental macro-molecular complexes of bacteria, and smaller ones, which display a remarkable exponential distribution of sizes. This distribution is "universal" in two respects: it holds for vastly different genomes, and for functionally distinct genes. Similar statistical laws have been reported previously in studies of bacterial genomes, and generally attributed to purifying selection or neutral processes. Here, we perform a new analysis based on the concept of parsimony, and find that the prevailing evolutionary mechanism behind the formation of small syntons is a selective process of gene aggregation. Altogether, our results imply a common evolutionary process that selectively shapes the organization and diversity of bacterial genomes.
研究动机与目标
- 为解决在系统发育偏差和采样不均的情况下,从多个细菌基因组中可靠推断显著的基因共定位问题。
- 定义更高阶的基因组单元——syntons——超越成对共线性,捕捉相互邻近的基因簇,这些基因簇可扩展至操纵子。
- 确定小syntons的观测大小分布是否源于中性过程,还是选择性进化机制。
- 厘清祖先遗传与选择性基因聚集在塑造共线性进化中各自的相对贡献。
提出的方法
- 以1,108个完整的细菌基因组和4,467个COGs(直系同源基因簇)为输入,利用共线性作为引导,对直系同源关系进行迭代优化。
- 应用系统发育距离阈值δ以定义有效的基因组采样,通过将基因组权重与其相似性成反比,减少近缘或过度代表物种带来的偏差。
- 使用均匀随机基因定位的零模型来定义基因邻近性的统计显著性,计算邻近性p值,并将其转换为指数分布的检验统计量。
- 将syntons定义为在多个基因组中相互邻近的基因最大集合,采用20 kb内的上下文邻近性标准。
- 采用基于简约性的模型推断共线性的进化动态,从基因组三元组和团簇中估计基因聚集(pA)和解体(pD)的概率。
- 使用上下文差异δ和序列差异δs来衡量基因组相似性,并标定基因组对之间的进化距离。
实验结果
研究问题
- RQ1如何在大规模、系统发育偏差显著的细菌基因组集合中,可靠地推断基因的显著共定位?
- RQ2在成对关系之外,保守基因邻近性的组织结构是怎样的?syntons与已知的功能单元(如操纵子)有何关系?
- RQ3在祖先遗传与选择性基因聚集之间,哪种进化机制最能解释小syntons的观测指数大小分布?
- RQ4小syntons的指数大小分布是否为跨多样化细菌基因组及功能类别的普遍特性?
主要发现
- 小syntons(通常不对应于已知的功能单元如操纵子)在多样化细菌基因组中表现出普遍的指数大小分布。
- 小syntons的指数大小分布对不同错误发现率具有鲁棒性,并在COG标签随机置换后消失,表明其并非数据结构的人为产物。
- 对于包含三个或更多基因的团簇,基因聚集概率(pA)高于解体概率(pD),强有力地表明选择性聚集驱动了共线性进化。
- 大syntons(对应于基本的大分子复合物)可能通过从共同祖先基因组中不完全解体而得以保留。
- 指数大小分布不能仅由中性过程或纯化选择解释,而更可能源于偏好基因聚类的选择性机制。
- 该方法成功识别出两类syntons:大而功能一致的单元,以及小而呈指数分布的单元,后者由选择性聚集所塑造。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。