[论文解读] Persistent spectral based machine learning (PerSpect ML) for drug design
该论文提出了一种名为持久谱机器学习(PerSpect ML)的新框架,将持久同调与谱图理论、单纯复形理论及超图理论相结合,为蛋白质-配体结合亲和力预测生成多尺度拓扑特征。通过在距离或相互作用强度阈值上应用过滤过程,从拉普拉斯特征值中提取持久谱变量(如持久重数、均值和能量),PerSpect ML 在 PDBbind-2007、PDBbind-2013 和 PDBbind-2016 数据集上均取得了最先进性能,皮尔逊相关系数(PCC)最高达 0.840,均方根误差(RMSE)低至 1.724 kcal/mol。
In this paper, we propose persistent spectral based machine learning (PerSpect ML) models for drug design. Persistent spectral models, including persistent spectral graph, persistent spectral simplicial complex and persistent spectral hypergraph, are proposed based on spectral graph theory, spectral simplicial complex theory and spectral hypergraph theory, respectively. Different from all previous spectral models, a filtration process, as proposed in persistent homology, is introduced to generate multiscale spectral models. More specifically, from the filtration process, a series of nested topological representations, i,e., graphs, simplicial complexes, and hypergraphs, can be systematically generated and their spectral information can be obtained. Persistent spectral variables are defined as the function of spectral variables over the filtration value. Mathematically, persistent multiplicity (of zero eigenvalues) is exactly the persistent Betti number (or Betti curve). We consider 11 persistent spectral variables and use them as the feature for machine learning models in protein-ligand binding affinity prediction. We systematically test our models on three most commonly-used databases, including PDBbind-2007, PDBbind-2013 and PDBbind-2016. Our results, for all these databases, are better than all existing models, as far as we know. This demonstrates the great power of our PerSpect ML in molecular data analysis and drug design.
研究动机与目标
- 开发一种新型机器学习框架,以捕捉生物分子的多尺度拓扑与谱特征,从而提升药物设计效果。
- 通过引入过滤过程,生成持久谱表示,以解决传统谱模型的局限性。
- 利用源自持久谱理论的拓扑、几何与代数不变量,增强蛋白质-配体结合亲和力预测中的特征表示。
- 在标准基准数据库(PDBbind-2007、-2013、-2016)上系统评估所提出的 PerSpect ML 模型,并与现有最先进模型进行比较。
- 证明持久谱变量在化学信息学与生物信息学应用中显著提升预测性能。
提出的方法
- 基于图、单纯复形与超图的谱理论,提出三种持久谱模型:持久谱图、持久谱单纯复形与持久谱超图。
- 引入一种过滤过程——受持久同调启发——系统地在递增的过滤值下生成拓扑表示(图、单纯复形、超图)。
- 定义了 11 种持久谱变量作为特征,包括持久重数(与贝蒂曲线相关)、持久均值、标准差、最大值、最小值、拉普拉斯图能量、广义均值能量、二阶谱矩、拟威纳指数与生成树数量。
- 使用距离(0.00–25.00 Å,步长 0.10 Å)和相互作用强度(0.00–1.00,步长 0.01)作为过滤参数,分别生成每套系统 250 个和 100 个拉普拉斯矩阵。
- 采用梯度提升树(GBT)模型,设置 40,000 个估计器、最大深度 6、学习率 0.001,并采用 70% 的子采样策略,以处理高维特征向量并防止过拟合。
- 结合两种模型的特征:基于原子间距离的 ES-IDM(36 种原子类型组合)与基于相互作用能的 ES-IEM(50 种原子类型组合),分别生成每分子 99,000 个与 55,000 个特征。
实验结果
研究问题
- RQ1持久谱理论能否被扩展至图、单纯复形与超图,以生成生物分子的多尺度拓扑表示?
- RQ2从拉普拉斯矩阵特征值经过滤过程生成的持久谱变量,能否捕捉与药物设计相关的生物结构与能量特征?
- RQ3与现有最先进模型相比,PerSpect ML 特征能否提升蛋白质-配体结合亲和力预测的准确性?
- RQ4不同过滤参数(距离 vs. 相互作用强度)如何影响 PerSpect ML 在结合亲和力预测中的性能?
- RQ5在药物发现的机器学习模型中,整合多种持久谱特征(如重数、能量、矩)是否比单一特征更有效?
主要发现
- PerSpect ML 在 PDBbind-2016 数据集上实现了最高的皮尔逊相关系数(PCC)0.840,优于所有先前报告的模型。
- 在 PDBbind-2016 上,模型实现了 1.724 kcal/mol 的均方根误差(RMSE),为所有测试现有模型中的最低值。
- 在 PDBbind-2007 上,模型实现了 PCC 0.829 和 RMSE 1.868 kcal/mol,均优于以往最先进结果。
- ES-IDM 与 ES-IEM 模型的组合在所有三个数据集上均优于单一模型,其中在 PDBbind-2007 上 PCC 为 0.836,在 PDBbind-2016 上为 0.840。
- 采用 10 次独立回归与中位数报告的梯度提升树模型,确保了性能指标的稳健性并降低了方差。
- 持久重数在数学上被证明与持久贝蒂数完全对应,验证了该框架的拓扑一致性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。