[论文解读] The evaluation of protein folding rate constant is improved by predicting the folding kinetic order with a SVM-based method
本研究提出一种基于SVM的方法(SVM-KO),用于预测蛋白质折叠动力学顺序(两态与多态)并提升折叠速率常数的预测精度。该方法通过将预测分离为两个独立的回归模型——分别针对每种动力学类别——实现改进。仅使用序列长度和接触序数,该方法在两态蛋白质上的相关系数为0.84(标准误SE = 0.90),在多态蛋白质上为0.79,显著优于单一模型的回归方法。
Protein folding is a problem of large interest since it concerns the mechanism by which the genetic information is translated into proteins with well defined three-dimensional (3D) structures and functions. Recently theoretical models have been developed to predict the protein folding rate considering the relationships of the process with tolopological parameters derived from the native (atomic-solved) protein structures. Previous works classified proteins in two different groups exhibiting either a single-exponential or a multi-exponential folding kinetics. It is well known that these two classes of proteins are related to different protein structural features. The increasing number of available experimental kinetic data allows the application to the problem of a machine learning approach, in order to predict the kinetic order of the folding process starting from the experimental data so far collected. This information can be used to improve the prediction of the folding rate. In this work first we describe a support vector machine-based method (SVM-KO) to predict for a given protein the kinetic order of the folding process. Using this method we can classify correctly 78% of the folding mechanisms over a set of 63 experimental data. Secondly we focus on the prediction of the logarithm of the folding rate. This value can be obtained as a linear regression task with a SVM-based method. In this paper we show that linear correlation of the predicted with experimental data can improve when the regression task is computed over two different sets, instead of one, each of them composed by the proteins with a correctly predicted two state or multistate kinetic order.
研究动机与目标
- 通过在回归模型中引入动力学顺序分类,提升蛋白质折叠速率常数的预测精度。
- 开发一种机器学习方法,仅利用结构参数预测蛋白质是否通过两态或多态机制折叠。
- 评估根据折叠动力学机制分离折叠速率预测是否能提升与实验数据的相关性。
- 通过采用不同序列间隔阈值的接触序数,评估局部作用与非局部作用在决定折叠动力学中的相对重要性。
- 提供一种可推广的、经交叉验证的框架,仅基于最少的结构输入预测折叠动力学。
提出的方法
- 使用63个经实验表征的单结构域蛋白,基于序列长度和接触序数,训练支持向量机(SVM)分类器以预测动力学顺序(两态或多态)。
- 通过可变截断半径(最优值为9 Å)和序列间隔阈值(最优值为≥6个残基)计算接触序数(CO),以区分局部与非局部相互作用。
- 采用10折交叉验证,确保SVM-KO模型的稳健性与泛化能力。
- 应用线性SVM回归,基于相同输入特征预测折叠速率的对数(log kf)。
- 通过将数据集根据正确预测的动力学顺序划分为两个子集(34个两态蛋白,15个多态蛋白),分别对每个子集执行独立的线性回归,从而提升预测性能。
- 使用相关系数(r)、标准误(SE)和马修斯相关系数(MCC)评估性能。
实验结果
研究问题
- RQ1机器学习模型能否基于序列长度和接触序数准确预测蛋白质是否遵循两态或多态折叠机制?
- RQ2与单一模型相比,将折叠速率预测划分为两个独立模型(分别针对两态和多态蛋白)是否能提升与实验数据的相关性?
- RQ3在预测折叠动力学时,接触序数计算的最优截断半径与序列间隔阈值是什么?
- RQ4局部作用与非局部作用如何影响折叠动力学顺序与速率常数的预测?
- RQ5将动力学顺序预测纳入模型,能在多大程度上提升折叠速率常数估计的准确性?
主要发现
- SVM-KO方法正确将78%的蛋白质分类为两态或多态折叠机制,马修斯相关系数为0.53。
- 当仅考虑可靠性指数为3的预测时,准确率提升至85%,在75%的数据集上相关系数达0.66。
- 最优接触序数计算采用9 Å的截断半径和序列间隔≥6,表明非局部相互作用对折叠动力学更具预测力。
- 针对两态和多态蛋白分别建立的线性回归模型,相关系数分别为0.84(SE = 0.90)和0.79(SE = 0.90),显著优于单一模型方法(r = 0.65,SE = 1.35)。
- 两个子集的平均标准误降低至约0.90,表明在分离动力学机制后,速率预测的精度得到提升。
- 结果证实,蛋白质长度和接触序数是决定折叠速率的关键因素,但其相对贡献在两态与多态折叠机制中有所不同。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。