Skip to main content
QUICK REVIEW

[论文解读] A three-state prediction of single point mutations on protein stability changes

Emidio Capriotti, Piero Fariselli|ArXiv.org|May 10, 2007
Protein Structure and Dynamics参考文献 17被引用 5
一句话总结

本文提出了一种三分类支持向量机预测器,根据序列或结构信息将蛋白质单一位点突变分类为稳定化(ΔΔG > 0.5 kcal/mol)、去稳定化(ΔΔG < -0.5 kcal/mol)或中性(|ΔΔG| ≤ 0.5 kcal/mol)。该方法基于序列的预测准确率为52%,基于结构的预测准确率为58%,显著优于随机预测,相关系数分别为0.30和0.39。

ABSTRACT

A basic question of protein structural studies is to which extent mutations affect the stability. This question may be addressed starting from sequence and/or from structure. In proteomics and genomics studies prediction of protein stability free energy change (DDG) upon single point mutation may also help the annotation process. The experimental SSG values are affected by uncertainty as measured by standard deviations. Most of the DDG values are nearly zero (about 32% of the DDG data set ranges from -0.5 to 0.5 Kcal/mol) and both the value and sign of DDG may be either positive or negative for the same mutation blurring the relationship among mutations and expected DDG value. In order to overcome this problem we describe a new predictor that discriminates between 3 mutation classes: destabilizing mutations (DDG0.5 Kcal/mol) and neutral mutations (-0.5&lt;=DDG&lt;=0.5 Kcal/mol). In this paper a support vector machine starting from the protein sequence or structure discriminates between stabilizing, destabilizing and neutral mutations. We rank all the possible substitutions according to a three state classification system and show that the overall accuracy of our predictor is as high as 52% when performed starting from sequence information and 58% when the protein structure is available, with a mean value correlation coefficient of 0.30 and 0.39, respectively. These values are about 20 points per cent higher than those of a random predictor.

研究动机与目标

  • 为解决实验测得的ΔΔG值中存在的模糊性问题,即许多突变导致变化接近零,从而模糊了突变与稳定性效应之间的关系。
  • 通过将突变重新分类为三个明确类别而非将ΔΔG视为连续变量,以提高预测准确性。
  • 开发一种鲁棒的预测器,利用蛋白质序列或三维结构来分类突变对蛋白质稳定性的影响。
  • 通过聚焦于生物上具有意义的突变类别,减少ΔΔG测量中的实验不确定性带来的噪声。

提出的方法

  • 训练支持向量机(SVM)将突变分为三类:稳定化、中性或去稳定化,分类阈值为±0.5 kcal/mol的ΔΔG。
  • 输入特征来自氨基酸序列或蛋白质三维结构,捕捉局部和全局结构背景信息。
  • 模型采用多分类框架,将每个氨基酸替换分配至三个稳定性结果类别之一。
  • 通过准确率和预测值与实验ΔΔG值之间的皮尔逊相关系数评估性能。
  • 该方法在包含实验测得ΔΔG值的单一位点突变数据集上进行测试,考虑了测量不确定性的影响。

实验结果

研究问题

  • RQ1与连续ΔΔG预测相比,三分类分类系统是否能提升对单一位点突变导致的蛋白质稳定性变化的预测能力?
  • RQ2引入三维结构信息在多大程度上提升了突变稳定性预测的准确性?
  • RQ3ΔΔG值的实验不确定性在多大程度上掩盖了突变与稳定性效应之间的关系?
  • RQ4基于序列或结构训练的机器学习模型能否可靠地区分稳定化、中性和去稳定化突变?

主要发现

  • 仅使用蛋白质序列信息时,预测器准确率达到52%,显著高于随机预测。
  • 在获得三维蛋白质结构信息后,准确率提升至58%,表明结构背景信息具有附加价值。
  • 序列预测的预测值与实验ΔΔG值之间的相关系数达到0.30,结构预测的相关系数达到0.39。
  • 三分类分类系统有效减少了实验不确定性带来的噪声,其中约32%的ΔΔG值落在中性范围(-0.5至0.5 kcal/mol)。
  • 该方法在准确率上比随机预测高出约20个百分点,表明具备有意义的预测能力。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。