[论文解读] 3D Protein Structure Predicted from Sequence
本文提出一种从氨基酸序列预测三维蛋白质结构的从头方法,利用进化共变性数据。通过将数据约束的最大熵模型应用于多序列比对,识别出统计上耦合的残基对(EICs),作为结构预测的远距离约束,无需同源建模或已知结构模板,即可在多种蛋白质家族中实现低Cα-RMSD误差(2.7–5.1 Å)。
The evolutionary trajectory of a protein through sequence space is constrained by function and three-dimensional (3D) structure. Residues in spatial proximity tend to co-evolve, yet attempts to invert the evolutionary record to identify these constraints and use them to computationally fold proteins have so far been unsuccessful. Here, we show that co-variation of residue pairs, observed in a large protein family, provides sufficient information to determine 3D protein structure. Using a data-constrained maximum entropy model of the multiple sequence alignment, we identify pairs of statistically coupled residue positions which are expected to be close in the protein fold, termed contacts inferred from evolutionary information (EICs). To assess the amount of information about the protein fold contained in these coupled pairs, we evaluate the accuracy of predicted 3D structures for proteins of 50-260 residues, from 15 diverse protein families, including a G-protein coupled receptor. These structure predictions are de novo, i.e., they do not use homology modeling or sequence-similar fragments from known structures. The resulting low Cα-RMSD error range of 2.7-5.1Å, over at least 75% of the protein, indicates the potential for predicting essentially correct 3D structures for the thousands of protein families that have no known structure, provided they include a sufficiently large number of divergent sample sequences. With the current enormous growth in sequence information based on new sequencing technology, this opens the door to a comprehensive survey of protein 3D structures, including many not currently accessible to the experimental methods of structural genomics. This advance has potential applications in many biological contexts, such as synthetic biology, identification of functional sites in proteins and interpretation of the functional impact of genetic variants.
研究动机与目标
- 开发一种仅基于序列数据和进化信息的从头三维蛋白质结构预测方法。
- 通过多序列比对中的统计共变性分析,识别折叠蛋白中空间上邻近的残基对。
- 评估在共变残基对中捕获的进化约束是否足以重建准确的三维结构。
- 为无已知实验结构的蛋白质家族(尤其是缺乏同源模板的蛋白质)提供结构预测能力。
- 提供一种可扩展的框架,利用迅速增长的序列数据对整个蛋白质组进行大规模结构调查。
提出的方法
- 构建一个基于多序列比对中观察到的残基对频率的约束最大熵模型。
- 使用反向伊辛模型(Potts模型)推断成对相互作用,识别出统计上耦合的残基对——称为进化信息接触(EICs)。
- 在三维结构预测流程中应用推断出的EICs作为远距离距离约束。
- 使用基于马尔可夫链蒙特卡洛的采样方法,生成与EIC约束和序列偏好一致的三维结构。
- 在无需已知模板、同源建模或片段库的情况下实现从头折叠。
- 使用实验测定的结构作为参照,通过Cα-均方根偏差(RMSD)验证预测结果。
实验结果
研究问题
- RQ1多序列比对中的进化共变性模式是否足以提供信息,以实现三维蛋白质结构的从头重建?
- RQ2推断出的残基接触(EICs)在多大程度上反映了折叠蛋白中真实的三维邻近性?
- RQ3该方法能否对无已知结构同源物的蛋白质实现准确的三维结构预测?
- RQ4该方法在不同蛋白质家族(包括G蛋白偶联受体等膜蛋白)中的结构预测精度如何变化?
- RQ5使用该方法实现可靠三维结构预测所需的最少分歧序列数量是多少?
主要发现
- 该方法成功实现了从头预测三维蛋白质结构,Cα-RMSD误差在至少75%的蛋白质骨架上介于2.7 Å至5.1 Å之间。
- 该方法在15种不同蛋白质家族中均实现了准确预测,包括一种G-蛋白偶联受体,证明其在具有挑战性的膜蛋白家族中的适用性。
- 预测结构完全基于进化共变性数据,未使用已知模板或片段库。
- 本研究表明,多序列比对中编码的进化约束包含足够信息,可用于推断三维折叠拓扑结构。
- 结果表明,只要有足够的序列多样性,该方法可实现对数千个未表征蛋白质家族的大规模、高通量三维结构预测。
- 该方法在不同大小(50–260个残基)和折叠类型的蛋白质中均表现稳健,表明其在未表征蛋白质组中的广泛适用性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。