[论文解读] Developing optimal nonlinear scoring function for protein design
本文提出了一种基于高斯核函数混合的非线性评分函数,通过区分天然蛋白序列与数百万种错误折叠序列,优化蛋白质序列设计。该方法在1400万种错误折叠序列中完美区分了440个天然蛋白,并在盲测中优于线性函数,证明非线性模型对于同时准确捕捉多个蛋白质的适应度景观至关重要。
Motivation. Protein design aims to identify sequences compatible with a given protein fold but incompatible to any alternative folds. To select the correct sequences and to guide the search process, a design scoring function is critically important. Such a scoring function should be able to characterize the global fitness landscape of many proteins simultaneously. Results. To find optimal design scoring functions, we introduce two geometric views and propose a formulation using mixture of nonlinear Gaussian kernel functions. We aim to solve a simplified protein sequence design problem. Our goal is to distinguish each native sequence for a major portion of representative protein structures from a large number of alternative decoy sequences, each a fragment from proteins of different fold. Our scoring function discriminate perfectly a set of 440 native proteins from 14 million sequence decoys. We show that no linear scoring function can succeed in this task. In a blind test of unrelated proteins, our scoring function misclassfies only 13 native proteins out of 194. This compares favorably with about 3-4 times more misclassifications when optimal linear functions reported in literature are used. We also discuss how to develop protein folding scoring function.
研究动机与目标
- 开发一种评分函数,能够在序列设计过程中同时建模多个蛋白质的适应度景观。
- 解决线性评分函数在区分结构多样的错误折叠序列与天然序列时的局限性。
- 通过捕捉超出成对残基接触的复杂非线性能量相互作用,提高蛋白质设计的准确性。
- 为构建适用于蛋白质设计与折叠的可推广评分函数框架提供支持。
提出的方法
- 将评分函数表述为非线性高斯核函数的混合,以建模蛋白质序列中复杂的非加和性相互作用。
- 利用序列空间的几何视角,定义基于核函数的优化问题,以分离天然序列与错误折叠序列。
- 应用凸优化框架,学习核权重以最大化天然序列与错误折叠序列之间的间隔。
- 使用包含440个天然蛋白和通过无间隙同源建模生成的1400万种序列错误折叠的训练集,学习最优评分函数。
- 通过在194个无关蛋白上的盲测验证方法,评估其泛化性能。
- 通过凸包分析证明,特征空间中不存在线性分隔器,因此必须采用非线性建模。
实验结果
研究问题
- RQ1非线性评分函数是否能在区分大量结构多样的错误折叠序列池中的天然蛋白序列方面优于线性函数?
- RQ2是否可能使用单一优化评分函数同时建模多个蛋白质的适应度景观?
- RQ3评分函数必须具备哪些几何与功能特性,才能确保对天然序列与错误折叠序列实现完美区分?
- RQ4为何线性评分函数在此蛋白质设计任务中失败?其失败的数学条件是什么?
- RQ5基于核函数的模型能否在训练中未见的无关蛋白上实现有效泛化?
主要发现
- 所提出的非线性评分函数在1400万种通过无间隙同源建模生成的序列错误折叠中,完美区分了440个天然蛋白序列。
- 不存在任何线性评分函数能在此任务中实现完美区分,其证明为:天然与错误折叠接触图差异向量的凸包包含原点。
- 在194个无关蛋白的盲测中,非线性函数仅错误分类13个天然序列,而文献中最佳线性函数错误分类数量约为其3至4倍。
- 该方法成功同时建模了多个蛋白质的全局适应度景观,表明其在蛋白质设计中具有普遍应用潜力。
- 基于核函数的学习方法有效捕捉了线性模型固有缺失的复杂非加和性相互作用。
- 该框架具有可推广性,可扩展以包含高阶相互作用及替代蛋白质表征方式,如显式氢键或残基描述符。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。