Skip to main content
QUICK REVIEW

[论文解读] Benchmarking structure-based three-dimensional molecular generative models using GenBench3D: ligand conformation quality matters

Benoît Baillif, Jason C. Cole|arXiv (Cornell University)|Jul 5, 2024
Genetics, Bioinformatics, and Biomedical Research被引用 4
一句话总结

该论文提出了GenBench3D,一个用于基于结构的3D分子生成模型的基准测试,通过新颖的Validity3D度量标准评估配体构象质量,该度量标准将键长和价键角与剑桥结构数据库的参考值进行对比。仅有0–11%的生成分子具有有效构象,但局部松弛使Validity3D得分提升了至少40%,揭示出原始生成的分子往往高估了结合亲和力,尤其是在Vina评分中表现明显。

ABSTRACT

Three-dimensional (3D) deep molecular generative models offer the advantage of goal-directed generation based on 3D-dependent properties, such as binding affinity for structure-based design within binding pockets. Traditional benchmarks created to evaluate SMILES or molecular graphs generators, such as GuacaMol or MOSES, are limited to evaluate 3D generators as they do not assess the quality of the generated molecular conformation. In this work, we hence developed GenBench3D, which implements a new benchmark for models producing molecules within a binding pocket. Our main contribution is the Validity3D metric, evaluating the conformation quality using the likelihood of bond lengths and valence angles based on reference values observed in the Cambridge Structural Database. The LiGAN, 3D-SBDD, Pocket2Mol, TargetDiff, DiffSBDD and ResGen models were benchmarked. We show that only between 0% and 11% of generated molecules have valid conformations. Performing local relaxation of generated molecules in the pocket considerably improved the Validity3D for all models by a minimum increase of 40%. For LiGAN, 3D-SBDD, or TargetDiff, the set of valid relaxed molecules shows on average higher Vina score (i.e. worse) than the set of raw generated molecules, indicating that the binding affinity of raw generated molecules might be overestimated. Using the other scoring functions, that give higher importance to ligand strain, only yield improved scores when using valid relaxed molecules. Using valid relaxed molecules, TargetDiff and Pocket2Mol show better median Vina, Glide and Gold PLP scores than other models. We have publicly released GenBench3D on GitHub for broader use: https://github.com/bbaillif/genbench3d

研究动机与目标

  • 为解决缺乏评估基于结构的生成模型中3D分子构象质量的基准测试的问题。
  • 识别并量化在结合口袋中生成分子的几何无效构象的普遍性。
  • 评估构象松弛是否能同时提升几何有效性与结合亲和力预测的准确性。
  • 通过标准化基准测试,比较六种3D生成模型——LiGAN、3D-SBDD、Pocket2Mol、TargetDiff、DiffSBDD和ResGen的性能。
  • 公开发布GenBench3D,以支持未来3D分子生成模型的可复现评估。

提出的方法

  • 开发了GenBench3D,一个用于在结合口袋背景下评估3D分子生成模型的基准测试框架。
  • 提出了Validity3D度量标准,通过计算键长和价键角相对于剑桥结构数据库参考值的似然性来评估几何质量。
  • 在预定义的结合口袋内,使用六种最先进的3D生成模型生成分子。
  • 采用分子力学方法(如OPLS4力场)进行局部构象松弛,以提升几何质量。
  • 使用多种打分函数评估结合亲和力:AutoDock Vina、Glide和Gold PLP。
  • 在几何有效性与打分函数性能方面,对比原始生成分子与其松弛后版本的表现。

实验结果

研究问题

  • RQ1当置于结合口袋中时,3D生成模型生成的分子中有多少比例具有几何上有效的构象?
  • RQ2局部构象松弛在多大程度上改善了生成分子的几何有效性?
  • RQ3由于构象质量较差,原始生成分子的结合亲和力是否往往被高估?
  • RQ4在松弛处理后,哪种3D生成模型能产生几何上最有效且打分最高的分子?
  • RQ5不同打分函数(Vina、Glide、Gold PLP)在构象松弛后,其结合亲和力预测准确性的表现如何?

主要发现

  • 在使用Validity3D度量标准评估时,六种模型生成的分子中仅有0%至11%在松弛前具有有效构象。
  • 所有模型中,局部构象松弛使Validity3D得分至少提升了40%,部分模型的提升幅度显著更高。
  • 对于LiGAN、3D-SBDD和TargetDiff,松弛后分子的Vina得分平均低于原始生成分子,表明原始打分可能高估了结合亲和力。
  • 在使用有效且已松弛的分子时,TargetDiff和Pocket2Mol在Vina、Glide和Gold PLP打分函数中的中位数得分优于其他模型。
  • 对配体应变进行惩罚的打分函数(如Glide、Gold PLP)仅在应用于已松弛且有效的构象时才表现出性能提升。
  • GenBench3D已发布于GitHub,以支持未来3D分子生成模型的可复现基准测试。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。