Skip to main content
QUICK REVIEW

[论文解读] On Design of Optimal Nonlinear Kernel Potential Function for Protein Folding and Protein Design

Changyu Hu, Xiang Li|arXiv (Cornell University)|Jan 31, 2003
Protein Structure and Dynamics参考文献 58被引用 3
一句话总结

本文提出一种基于高斯混合模型的非线性核势函数,以改进蛋白质折叠与序列设计,其性能优于传统的线性接触势函数。通过二次规划最小化泛化误差界,该方法在1400万个无缺口拖带构象中实现了对440个天然蛋白质及其序列的完美区分,相较于线性势函数,在独立测试中显著降低了误分类率。

ABSTRACT

Potential functions are critical for computational studies of protein structure prediction, folding, and sequence design. A class of widely used potentials for coarse grained models of proteins are contact potentials in the form of weighted linear sum of pairwise contacts. However, these potentials have been shown to be unsuitable choices because they cannot stabilize native proteins against a large number of decoys generated by gapless threading. We develop an alternative framework for designing protein potential. We describe how finding optimal protein potential can be understood from two geometric viewpoints, and we derive nonlinear potentials using mixture of Gaussian kernel functions for folding and design. The optimization criterion for obtaining parameters of the potential is to minimize bounds on the generalization error of discriminating protein structures and decoys not used in training. In our experiment we use a training set of 440 protein structures repre senting a major portion of all known protein structures, and about 14 million structure decoys and sequence decoys obtained by gapless threading. We succeeded in obtaining nonlinear potential with perfect discrimination of the 440 native structures and native sequences. For the more challenging task of sequence design when decoys are obtained by gapless threading, we show that there is no linear potential with perfect discrimination of all 440 native sequences. Results on an independent test set of 194 proteins also showed that nonlinear kernel potential performs well.

研究动机与目标

  • 解决线性接触势函数在对抗大规模构象集合时对天然蛋白质结构稳定性不足的问题。
  • 开发一种更复杂的势函数形式,以捕捉蛋白质能量景观中的复杂非线性相互作用。
  • 在蛋白质折叠(结构区分)与蛋白质设计(序列相容性)任务中均提升性能。
  • 证明非线性核势函数可在传统线性势函数失效的情况下实现完美分类,尤其在具有挑战性的构象条件下。
  • 提供一种可推广的势函数设计框架,适用于多种蛋白质表征与相互作用形式。

提出的方法

  • 将蛋白质势函数表述为高斯核函数的混合,以实现残基-残基相互作用的非线性建模。
  • 使用二次规划优化势函数参数,通过最小化结构与序列区分的理论泛化误差界。
  • 采用无缺口拖带法生成1400万个结构与序列构象,用于训练与验证。
  • 从几何视角解释优化问题,将其与能量差值的凸包分离相联系。
  • 在440个蛋白质的训练集与194个蛋白质的独立测试集上验证该方法,并与最优线性势函数进行比较。
  • 通过能量差值向量的凸包分析,推导出实现完美区分的理论条件。

实验结果

研究问题

  • RQ1非线性核势函数能否在区分天然蛋白质结构与构象方面优于标准线性接触势函数?
  • RQ2是否可能在传统线性势函数失效的情况下,利用非线性势函数实现对无缺口拖带构象中天然序列的完美分类?
  • RQ3在独立测试集中,非线性核势函数在折叠与设计任务中的性能相较于线性势函数如何?
  • RQ4当天然结构数量超过300个时,势函数的形式在实现有效区分中起到何种作用?
  • RQ5基于泛化误差界优化的框架能否产生更鲁棒、更具泛化能力的蛋白质势函数?

主要发现

  • 非线性核势函数在1400万个无缺口拖带构象中实现了对全部440个天然蛋白质结构与序列的完美区分,而线性势函数在此任务中失败。
  • 在194个蛋白质的独立测试集中,非线性势函数仅误分类3个结构与14个序列,而最优线性势函数则误分类7个结构与37个序列。
  • 该方法将序列设计中的误分类率降低至最优线性势函数的40%,在具有挑战性的场景中表现出显著改进。
  • 理论分析表明,只有当能量差值向量的凸包不包含原点时,才可能实现完美区分,而该非线性核形式恰好满足此条件。
  • 非线性势函数在折叠与设计任务中均优于线性与统计势函数,尤其在通过无缺口拖带生成构象时表现更优。
  • 结果表明,仅依赖接触加权线性和的复杂函数形式不足以实现有效的蛋白质能量函数设计,必须引入更复杂的表达形式。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。