Skip to main content
QUICK REVIEW

[论文解读] Screening of Therapeutic Agents for COVID-19 using Machine Learning and Ensemble Docking Simulations

Rohit Batra, Henry Chan|arXiv (Cornell University)|Apr 8, 2020
Computational Drug Discovery Methods参考文献 26被引用 11
一句话总结

本研究提出了一种混合机器学习与集合对接方法,通过预测配体与SARS-CoV-2刺突蛋白及其与人ACE2复合物的结合亲和力,实现对COVID-19治疗药物的快速筛选。该方法从一百万分子的数据库中识别出19,000个高潜力候选分子,显著加速了虚拟筛选过程,同时通过基于高保真对接数据训练的验证随机森林模型保持了高准确性。

ABSTRACT

The world has witnessed unprecedented human and economic loss from the COVID-19 disease, caused by the novel coronavirus SARS-CoV-2. Extensive research is being conducted across the globe to identify therapeutic agents against the SARS-CoV-2. Here, we use a powerful and efficient computational strategy by combining machine learning (ML) based models and high-fidelity ensemble docking simulations to enable rapid screening of possible therapeutic molecules (or ligands). Our screening is based on the binding affinity to either the isolated SARS-CoV-2 S-protein at its host receptor region or to the Sprotein-human ACE2 interface complex, thereby potentially limiting and/or disrupting the host-virus interactions. We first apply our screening strategy to two drug datasets (CureFFI and DrugCentral) to identify hundreds of ligands that bind strongly to the aforementioned two systems. Candidate ligands were then validated by all atom docking simulations. The validated ML models were subsequently used to screen a large bio-molecule dataset (with nearly a million entries) to provide a rank-ordered list of ~19,000 potentially useful compounds for further validation. Overall, this work not only expands our knowledge of small-molecule treatment against COVID-19, but also provides an efficient pathway to perform high-throughput computational drug screening by combining quick ML surrogate models with expensive high-fidelity simulations, for accelerating the therapeutic cure of diseases.

研究动机与目标

  • 通过克服高保真对接模拟的计算瓶颈,加速针对SARS-CoV-2的虚拟筛选。
  • 开发准确的机器学习代理模型,用于预测靶向S蛋白及S蛋白:ACE2界面的配体结合亲和力(Vina得分)。
  • 将筛选范围扩展至已知药物之外的广阔化学空间(约100万个生物分子),并利用经过验证的机器学习模型实现。
  • 识别出具有强结合亲和力和类药物性质的候选化合物,以供进一步实验验证。
  • 提供一种可扩展、高效的框架,适用于新兴疾病高通量计算药物发现。

提出的方法

  • 训练了两个独立的随机森林回归模型,用于预测配体与孤立S蛋白及S蛋白:ACE2复合物结合的AutoDock Vina得分。
  • 使用分层分子描述符,捕捉在多个长度尺度上的几何与化学特征,作为机器学习输入的配体表示。
  • 在来自CureFFI和DrugCentral的约5,500个配体数据集上训练模型,该数据集源自先前的高保真集合对接模拟。
  • 通过数百个高潜力候选分子的全原子对接模拟,对机器学习预测结果进行验证,确保准确性。
  • 将经过验证的机器学习模型应用于BindingDB数据库中约100万个生物分子的筛选,以识别高亲和力结合剂。
  • 采用基于阈值的筛选标准(Vina得分低于截止值)对结果进行排序和筛选,从BindingDB中选出约19,000个有前景的候选分子。

实验结果

研究问题

  • RQ1基于高保真对接数据训练的机器学习模型能否准确预测SARS-CoV-2治疗候选分子的结合亲和力?
  • RQ2机器学习代理模型在多大程度上可降低虚拟筛选的计算成本,同时保持准确性?
  • RQ3该机器学习-对接混合方法在探索已知药物之外的大型化学空间方面具有多大的可扩展性?
  • RQ4来自BindingDB数据集的高排名化合物中,哪些表现出强结合亲和力和有利的类药物性质?
  • RQ5该方法能否识别出具有潜在再利用价值的FDA批准药物或天然化合物用于COVID-19治疗?

主要发现

  • 机器学习模型表现出高预测准确性:在CureFFI和DrugCentral中筛选的187种FDA批准的配体中,有175种被预测为能与孤立S蛋白及S蛋白:ACE2复合物均强结合,该结果经对接模拟验证。
  • 与对187个候选分子进行对接模拟约需2天相比,该方法利用机器学习预测在一天内完成对约100万个生物分子的筛选。
  • 共从BindingDB数据集中识别出19,000个高亲和力配体,其结合亲和力阈值同时满足两个靶标系统的要求。
  • 优选候选分子包括Fidarestat、Quercetin、Myricetin、S-columbianetin、Indirubin和Cupressuflavone,均表现出强预测结合亲和力及相关的生物活性。
  • 许多高排名化合物,包括多种FDA批准的药物,符合Lipinski的“五规则”,表明其具有有利的类药物性质。
  • 该混合机器学习-对接工作流程提供了一条可扩展、高效且准确的高通量虚拟筛选管道,适用于SARS-CoV-2及其他疾病的治疗药物发现。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。