Skip to main content
QUICK REVIEW

[论文解读] Assessing interaction recovery of predicted protein-ligand poses

David Errington, Constantin Schneider|arXiv (Cornell University)|Sep 30, 2024
Cell Image Analysis TechniquesBiochemistry, Genetics and Molecular Biology被引用 3
一句话总结

本文引入蛋白质-配体相互作用指纹(PLIF)恢复作为评估蛋白质-配体构象预测模型的关键指标,表明尽管均方根偏差(RMSD)和PoseBuster有效性较低,许多基于机器学习的对接和共折叠模型仍无法重现关键相互作用,如氢键和卤素键。研究发现,经典对接方法在相互作用恢复方面显著优于机器学习方法,凸显了在模型训练中引入显式药效团感知损失函数的必要性。

ABSTRACT

The field of protein-ligand pose prediction has seen significant advances in recent years, with machine learning-based methods now being commonly used in lieu of classical docking methods or even to predict all-atom protein-ligand complex structures. Most contemporary studies focus on the accuracy and physical plausibility of ligand placement to determine pose quality, often neglecting a direct assessment of the interactions observed with the protein. In this work, we demonstrate that ignoring protein-ligand interaction fingerprints can lead to overestimation of model performance, most notably in recent protein-ligand cofolding models which often fail to recapitulate key interactions.

研究动机与目标

  • 解决过度依赖RMSD和PoseBuster有效性作为构象预测质量唯一指标的问题。
  • 证明相互作用恢复——尤其是氢键和卤素键等关键相互作用——是构象有效性的必要但常被忽视的标准。
  • 评估经典对接、机器学习对接和共折叠模型在恢复真实蛋白质-配体相互作用指纹(PLIF)方面的表现。
  • 倡导在机器学习模型中整合显式药效团或PLIF敏感的损失函数,以提升预测构象的生物学相关性。

提出的方法

  • 使用ProLIF软件包(v2.0.3)计算蛋白质-配体相互作用指纹(PLIF),重点关注特定相互作用类型:氢键、卤素键、π-π堆积、阳离子-π、π-阳离子和离子相互作用。
  • 应用自定义距离阈值:氢键为3.7 Å,阳离子-π为5.5 Å,离子相互作用为5.0 Å,其余参数均设为默认值。
  • PLIF恢复率定义为正确恢复的相互作用数与晶体结构中真实相互作用数的比值,各类相互作用的召回率分别计算。
  • 模型在PoseBusters测试集(308个复合物)上进行评估,RMSD和PoseBuster有效性作为基线指标,而PLIF恢复率作为主要的新指标。
  • 分析比较了经典对接(GOLD)、机器学习对接(DiffDock-L)和共折叠模型(RoseTTAFold-AllAtom、Umol等)在相互作用恢复和几何精度方面的表现。
  • 采用修改后的PoseBuster有效性定义,排除配体RMSD以隔离相互作用恢复对构象质量的贡献。
Figure 1: Left: Two-dimensional representation of the ligand EZO and its four interactions with the crystal structure 6M2B. Basic residues are shown in blue and residues containing a sulfur atom are shown in yellow. Right: Docked poses generated with GOLD, DiffDock-L and RosettaFold-AllAtom showing
Figure 1: Left: Two-dimensional representation of the ligand EZO and its four interactions with the crystal structure 6M2B. Basic residues are shown in blue and residues containing a sulfur atom are shown in yellow. Right: Docked poses generated with GOLD, DiffDock-L and RosettaFold-AllAtom showing

实验结果

研究问题

  • RQ1最先进基于机器学习的蛋白质-配体构象预测模型在多大程度上能恢复与晶体结构相比的关键蛋白质-配体相互作用?
  • RQ2PLIF恢复率与RMSD和PoseBuster有效性之间的相关性如何?它能否揭示标准指标所遗漏的性能局限?
  • RQ3为何共折叠模型尽管RMSD较低,却无法恢复生物相关性相互作用,如氢键和卤素键?
  • RQ4整合显式药效团或相互作用感知的损失函数是否能提升机器学习预测构象的生物学相关性?

主要发现

  • 经典对接(GOLD)在6M2B复合物中100%恢复了氢键及所有关键相互作用,而DiffDock-L遗漏了卤素键并改变了氢键模式。
  • 尽管RMSD仅为1.5 Å,DiffDock-L在6M2B复合物中仅恢复了75%的PLIF,表明其对关键相互作用的重现部分但不完整。
  • RoseTTAFold-AllAtom的RMSD为2.19 Å,但未能恢复任何真实相互作用,尽管其PoseBuster有效性成立。
  • 平均而言,机器学习对接和共折叠模型恢复的氢键和π-π堆积相互作用显著少于经典对接,后者在相互作用恢复方面始终优于机器学习方法。
  • 研究发现PLIF恢复率与RMSD相关性微弱,表明低RMSD并不能保证恢复生物相关性相互作用。
  • 共折叠模型常产生存在空间位阻和错误取向的构象,导致出现虚假的离子相互作用,取代了晶体结构中真实的阳离子-π或π-π堆积相互作用。
Figure 2: The ratio of predicted protein-ligand complex structures for each model passing checks on ligand positioning (RMSD $\leq$ 2Å), physicality (PoseBuster-valid) and interaction recovery (PLIF-valid).
Figure 2: The ratio of predicted protein-ligand complex structures for each model passing checks on ligand positioning (RMSD $\leq$ 2Å), physicality (PoseBuster-valid) and interaction recovery (PLIF-valid).

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。