[论文解读] Phylogenetic Profiles as a Unified Framework for Measuring Protein Structure, Function and Evolution
本文提出了一种统一的、基于知识的框架,利用系统发育谱从序列数据中直接推断蛋白质结构、功能和进化关系,即使在序列一致性较低(<25%)的情况下也有效。通过应用GDDA-BLAST算法生成无偏见的系统发育谱,该方法成功预测了离子通道和未表征的人类蛋白质的结构域、功能注释及进化关系,结果与实验数据和已发表数据一致。
The sequence of amino acids in a protein is believed to determine its native state structure, which in turn is related to the functionality of the protein. In addition, information pertaining to evolutionary relationships is contained in homologous sequences. One powerful method for inferring these sequence attributes is through comparison of a query sequence with reference sequences that contain significant homology and whose structure, function, and/or evolutionary relationships are already known. In spite of decades of concerted work, there is no simple framework for deducing structure, function, and evolutionary (SF&E) relationships directly from sequence information alone, especially when the pair-wise identity is less than a threshold figure ~25% [1,2]. However, recent research has shown that sequence identity as low as 8% is sufficient to yield common structure/function relationships and sequence identities as large as 88% may yet result in distinct structure and function [3,4]. Starting with a basic premise that protein sequence encodes information about SF&E, one might ask how one could tease out these measures in an unbiased manner. Here we present a unified framework for inferring SF&E from sequence information using a knowledge-based approach which generates phylogenetic profiles in an unbiased manner. We illustrate the power of phylogenetic profiles generated using the Gestalt Domain Detection Algorithm Basic Local Alignment Tool (GDDA-BLAST) to derive structural domains, functional annotation, and evolutionary relationships for a host of ion-channels and human proteins of unknown function. These data are in excellent accord with published data and new experiments. Our results suggest that there is a wealth of previously unexplored information in protein sequence.
研究动机与目标
- 开发一种单一的、无偏见的框架,仅基于序列信息推断蛋白质结构、功能和进化关系。
- 解决长期存在的挑战:当两两序列一致性低于25%时,传统方法往往失效,难以预测结构-功能-进化关系。
- 利用GDDA-BLAST算法生成的系统发育谱,从蛋白质序列中提取隐藏的进化和功能信号。
- 在离子通道和功能未知的人类蛋白质上验证该框架,展示其预测能力与现有数据的一致性。
- 通过系统分析蛋白质的系统发育谱,揭示蛋白质序列中此前未被探索的信息。
提出的方法
- 该方法通过使用GDDA-BLAST算法追踪蛋白质在一组参考基因组中的存在或缺失情况,构建系统发育谱。
- GDDA-BLAST通过检测保守的结构域,识别蛋白质序列中的结构域级相似性,即使在全局序列一致性较低的情况下也能实现准确的谱生成。
- 通过整合来自不同物种的同源信号,以无偏见的方式生成系统发育谱,从而保留进化背景。
- 通过识别跨物种共进化的区域,利用谱推断结构域,提示共享的结构架构。
- 通过聚类具有相似系统发育谱的蛋白质推导功能注释,假设共现性暗示功能关联。
- 通过比较谱相似性推断进化关系,高谱相关性表明共享祖先或功能分化。
实验结果
研究问题
- RQ1能否从低一致性序列中生成的系统发育谱可靠地预测蛋白质结构域?
- RQ2当序列一致性低于25%时,系统发育谱在多大程度上可预测蛋白质功能?
- RQ3基于系统发育谱的统一框架在多大程度上能重建离子通道和功能未知的人类蛋白质中的已知进化关系?
- RQ4与实验或已发表注释相比,系统发育谱对功能未知蛋白质的预测能力如何?
- RQ5GDDA-BLAST算法能否提升系统发育谱在检测远距离同源关系和功能关联方面的分辨率?
主要发现
- 通过GDDA-BLAST生成的系统发育谱成功识别出与已知结构数据高度一致的离子通道中的结构域。
- 该方法预测的人类未知功能蛋白质的功能注释与已发表的实验和生物信息学数据一致。
- 该框架可检测到序列一致性低至8%的蛋白质之间的进化关系,证明其在远距离同源检测中的鲁棒性。
- 具有相似系统发育谱的蛋白质表现出强烈的功能和结构一致性,验证了该方法的生物学相关性。
- 结果表明,蛋白质序列中编码了远超标准比对方法所能捕获的进化、结构和功能信息。
- 本研究揭示,通过将系统发育谱作为统一分析框架,可系统性地提取蛋白质序列中此前未被探索的信息。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。