Skip to main content
QUICK REVIEW

[论文解读] De novo PROTAC design using graph-based deep generative models

Divya Nori, Connor W. Coley|arXiv (Cornell University)|Nov 4, 2022
Protein Degradation and Inhibitors被引用 14
一句话总结

该论文提出了一种基于图的深度生成模型,通过策略梯度强化学习,利用提升树代理模型进行降解活性预测,从空图中从头设计新型PROTAC样分子。微调后,该模型将IRAK3的预测降解活性从50%提升至80%以上,且化学结构有效性接近完美,证明了对大型复杂PROTAC(最多140个重原子)的有效生成。

ABSTRACT

PROteolysis TArgeting Chimeras (PROTACs) are an emerging therapeutic modality for degrading a protein of interest (POI) by marking it for degradation by the proteasome. Recent developments in artificial intelligence (AI) suggest that deep generative models can assist with the de novo design of molecules with desired properties, and their application to PROTAC design remains largely unexplored. We show that a graph-based generative model can be used to propose novel PROTAC-like structures from empty graphs. Our model can be guided towards the generation of large molecules (30--140 heavy atoms) predicted to degrade a POI through policy-gradient reinforcement learning (RL). Rewards during RL are applied using a boosted tree surrogate model that predicts a molecule's degradation potential for each POI. Using this approach, we steer the generative model towards compounds with higher likelihoods of predicted degradation activity. Despite being trained on sparse public data, the generative model proposes molecules with substructures found in known degraders. After fine-tuning, predicted activity against a challenging POI increases from 50% to >80% with near-perfect chemical validity for sampled compounds, suggesting this is a promising approach for the optimization of large, PROTAC-like molecules for targeted protein degradation.

研究动机与目标

  • 开发一种用于PROTAC的从头分子生成框架,PROTAC是大型复杂双功能分子,可靶向蛋白降解。
  • 解决尽管公开训练数据稀疏,仍设计出具有高降解潜力PROTAC的挑战。
  • 实现无需预设子结构的原子级生成PROTAC样结构。
  • 通过强化学习提高生成分子中预测高降解活性分子的可能性。
  • 提供一个开源、可复现的计算工作流程,用于基于可解释机器学习和生成建模的PROTAC从头设计。

提出的方法

  • 使用基于图的深度生成模型(GraphINVENT)从空图出发,逐原子生成PROTAC样分子,学习化学价态和连接规则。
  • 在公开PROTAC数据上训练提升树代理模型,以标量分数预测降解活性(DC50),作为强化学习的奖励函数。
  • 应用策略梯度强化学习,引导生成模型向预测降解活性更高的分子发展,使用多目标评分函数。
  • 通过200步强化学习微调模型,同时优化预测活性和化学有效性。
  • 该方法应用于IRAK3靶点,生成无先前子结构约束的新型PROTAC样结构。
  • 对生成分子进行子结构相似性与化学有效性评估(最终模型为100%有效)。

实验结果

研究问题

  • RQ1基于图的深度生成模型能否在无种子子结构的情况下,从空图中从头设计新型PROTAC样分子?
  • RQ2使用降解活性代理模型的强化学习能否提高生成分子中预测活性PROTAC的比例?
  • RQ3在公开PROTAC数据稀疏且不均衡的条件下,模型在多大程度上能生成化学有效、大分子量(30–140个重原子)且含有已知降解剂子结构的分子?
  • RQ4在有限且不平衡的公开数据上训练时,提升树代理模型在预测PROTAC降解潜力(DC50)方面的表现如何?
  • RQ5在策略梯度强化学习中集成记忆感知损失函数,能否增强高活性PROTAC候选的生成?

主要发现

  • 基于图的深度生成模型成功生成了最多含139个重原子的PROTAC样分子,且最终微调模型的化学有效性达100%。
  • 经过200步强化学习后,IRAK3靶点的预测活性PROTAC比例从53%提升至82%。
  • 模型最终输出在具有挑战性的IRAK3靶点上实现了超过80%的预测降解活性率,显著高于初始50%基线。
  • 生成的顶级分子包含了在protac-db中常见于已知PROTAC降解剂的子结构,表明其具有生物学相关性。
  • 代理模型在DC50预测上的测试AUC达到0.87,但受限于公开数据稀疏且不均衡,泛化能力有限。
  • 生成分子100%为新颖结构,与训练化合物无结构重叠,证实了真正的从头设计能力。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。