Skip to main content
QUICK REVIEW

[论文解读] Meaningful machine learning models and machine-learned pharmacophores from fragment screening campaigns

Carl Poelking, Gianni Chessari|arXiv (Cornell University)|Mar 25, 2022
Computational Drug Discovery Methods被引用 6
一句话总结

本研究提出了一类基于片段筛选数据训练的可解释机器学习模型,通过引入经实验验证的‘真实漏检’来提升预测准确性,并采用归因框架识别出具有化学意义的相互作用模式。其核心贡献是一种经过验证的、具有物理可解释性的方法,可将模型预测映射至特定药效团特征,即使在模型性能欠佳时,仍与专家标注的分子相互作用表现出高度一致性。

ABSTRACT

Machine learning (ML) is widely used in drug discovery to train models that predict protein-ligand binding. These models are of great value to medicinal chemists, in particular if they provide case-specific insight into the physical interactions that drive the binding process. In this study we derive ML models from over 50 fragment-screening campaigns to introduce two important elements that we believe are absent in most -- if not all -- ML studies of this type reported to date: First, alongside the observed hits we use to train our models, we incorporate true misses and show that these experimentally validated negative data are of significant importance to the quality of the derived models. Second, we provide a physically interpretable and verifiable representation of what the ML model considers important for successful binding. This representation is derived from a straightforward attribution procedure that explains the prediction in terms of the (inter-)action of chemical environments. Critically, we validate the attribution outcome on a large scale against prior annotations made independently by expert molecular modellers. We find good agreement between the key molecular substructures proposed by the ML model and those assigned manually, even when the model's performance in discriminating hits from misses is far from perfect. By projecting the attribution onto predefined interaction prototypes (pharmacophores), we show that ML allows us to formulate simple rules for what drives fragment binding against a target automatically from screening data.

研究动机与目标

  • 为解决蛋白质-配体结合机器学习模型中缺乏可靠负样本的问题,通过引入片段筛选实验中经验证的‘真实漏检’数据。
  • 开发一种机器学习框架,为结合预测提供物理解释且可验证的说明,超越黑箱模型。
  • 通过将机器学习模型的归因结果与经验丰富的分子建模专家独立标注的药效团信息进行对比,验证模型的可解释性。
  • 直接从筛选数据中利用机器学习推导出简单、基于规则的药效团模型,实现从苗头化合物优化到先导化合物阶段的自动化洞察生成。
  • 证明即使在缺乏共价或非共价结合基序先验知识的情况下,归因场仍能准确定位与结合相关的化学环境。

提出的方法

  • 本研究采用基于原子环境相似性的核方法,通过 $G_{nl}Y_{lm}$ 描述符表示,计算探针片段与训练集中分子之间的成对核值。
  • 采用改进的积分梯度方法计算归因权重,定义为 $ z_a = \frac{1}{\text{norm}} \times \text{integral over } \alpha \text{ of } \partial Z_A / \partial k_{ab} $,其中 $ k_{ab} $ 衡量原子 a 与 b 之间的相似性。
  • 定义局部化归因场 $ \varphi(\hat{z}) $ 为围绕每个原子中心的归因权重的高斯核加权平均,使用参数 $ \alpha $ 的高斯核,实现对关键相互作用的空间定位。
  • 该方法引入基线输入 $ x' $,并使用 z-score 归一化的归因权重 $ \hat{z} $,以在不同分子环境中稳定和标准化场表示。
  • 该框架将归因结果投影到预定义的药效团原型上,将抽象的模型输出转化为可解释的、基于规则的结合决定因素洞察。
  • 通过将机器学习生成的归因场与 18 个结合位点上专家分子建模者手工整理的药效团标注进行对比,严格验证了模型的可解释性。

实验结果

研究问题

  • RQ1将经实验验证的‘真实漏检’纳入训练数据,是否能显著提升基于片段筛选数据训练的机器学习模型的性能与可靠性?
  • RQ2机器学习模型在蛋白质-配体结合预测中,能在多大程度上生成物理解释且可验证的预测说明?
  • RQ3机器学习生成的归因场在识别关键结合亚结构方面,与专家标注的药效团的吻合程度如何?
  • RQ4在缺乏结构知识先验的情况下,模型的归因机制是否能可靠地定位与结合相关的化学环境?
  • RQ5模型输出在多大程度上可直接从筛选数据中转化为简单、基于规则的药效团模型?

主要发现

  • 将经实验验证的‘真实漏检’纳入训练数据,显著提升了模型的质量与鲁棒性,相较于仅使用观测到的‘命中’数据训练的模型表现更优。
  • 机器学习生成的归因场在空间上准确地定位了与结合相关的特征,即使在未提供共价结合信息的情况下,也能正确识别出 SARS-CoV-2 主蛋白酶共价抑制剂中的乙酰基为关键结合基序。
  • 对于非共价片段命中,归因场呈现出弥散的模式,与缺乏单一主导结合基序的特性一致,解释了此类情况下模型预测准确性较低的原因。
  • 即使模型分类性能不完美,其归因识别出的关键分子亚结构与专家分子建模者手工分配的结构之间仍表现出高度一致性。
  • 当将归因场投影到预定义的药效团原型上时,可生成简单、可解释的片段结合规则,实现从筛选数据中自动化生成药效团假设。
  • 通过自助抽样法计算归因场的标准差,提供了可靠的不确定性度量,增强了对模型可解释性结果的信心。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。