[论文解读] Unveiling the molecular mechanism of SARS-CoV-2 main protease inhibition from 92 crystal structures
本研究提出了一种数学增强的深度学习(MathDL)框架,用于预测并排序92个SARS-CoV-2主蛋白酶(M¹⁷)抑制剂复合物的结合亲和力,基于X射线晶体结构,在经过筛选的SARS-CoV-2数据集上实现了0.751的皮尔逊相关系数。该方法识别出Gly143为最有利于形成氢键的位点,并揭示了45种共价抑制剂靶向Cys145,为COVID-19治疗药物的结构基础设计提供了支持。
Currently, there is no effective antiviral drugs nor vaccine for coronavirus disease 2019 (COVID-19) caused by acute respiratory syndrome coronavirus 2 (SARS-CoV-2). Due to its high conservativeness and low similarity with human genes, SARS-CoV-2 main protease (M$^{ ext{pro}}$) is one of the most favorable drug targets. However, the current understanding of the molecular mechanism of M$^{ ext{pro}}$ inhibition is limited by the lack of reliable binding affinity ranking and prediction of existing structures of M$^{ ext{pro}}$-inhibitor complexes. This work integrates mathematics and deep learning (MathDL) to provide a reliable ranking of the binding affinities of 92 SARS-CoV-2 M$^{ ext{pro}}$ inhibitor structures. We reveal that Gly143 residue in M$^{ ext{pro}}$ is the most attractive site to form hydrogen bonds, followed by Cys145, Glu166, and His163. We also identify 45 targeted covalent bonding inhibitors. Validation on the PDBbind v2016 core set benchmark shows the MathDL has achieved the top performance with Pearson's correlation coefficient ($R_p$) being 0.858. Most importantly, MathDL is validated on a carefully curated SARS-CoV-2 inhibitor dataset with the averaged $R_p$ as high as 0.751, which endows the reliability of the present binding affinity prediction. The present binding affinity ranking, interaction analysis, and fragment decomposition offer a foundation for future drug discovery efforts.
研究动机与目标
- 解决92个SARS-CoV-2 M¹⁷-抑制剂晶体结构缺乏实验结合亲和力数据的问题。
- 在实验数据稀疏的情况下,开发一种可靠的计算框架,用于对抑制剂结合亲和力进行排序。
- 识别出对M¹⁷高亲和力抑制至关重要的蛋白质残基和分子片段。
- 为基于结构的SARS-CoV-2靶向药物发现提供一个经过验证、数据驱动的资源。
提出的方法
- 在PDBbind v2019中的119个高质量M¹⁷-抑制剂复合物及一个经过筛选的SARS-CoV-2抑制剂集合上,训练了混合数学与深度学习(MathDL)模型。
- 利用PDBbind v2019中的17,382个通用蛋白-配体复合物对模型进行微调,并在PDBbind v2016核心集上进行验证。
- 采用多任务和单任务学习设置,以提高结合亲和力预测的泛化能力和鲁棒性。
- 使用MathPose为87个缺乏结构数据的抑制剂生成精确的3D构象,确保输入表示的一致性。
- 综合11个训练好的MathDL模型(5个单任务、5个多任务)的预测结果,以增强可靠性并降低方差。
- 应用BRICS算法对顶级抑制剂进行片段分解,以识别指导先导化合物优化的关键药效团结构特征。
实验结果
研究问题
- RQ1在无实验结合亲和力报告的92个SARS-CoV-2 M¹⁷-抑制剂复合物中,其相对结合亲和力的排序如何?
- RQ2M¹⁷中哪些氨基酸残基最常参与与抑制剂的氢键或共价相互作用?
- RQ3与现有评分函数相比,MathDL框架在预测结合亲和力方面的准确性如何?
- RQ4在高亲和力抑制剂中,哪些分子片段最为常见,且如何指导先导化合物设计?
- RQ5共价相互作用与非共价相互作用在M¹⁷抑制中对抑制剂效力的贡献分别是什么?
主要发现
- 在PDBbind v2016核心集上,MathDL模型实现了0.858的皮尔逊相关系数,优于现有评分函数。
- 在经过筛选的SARS-CoV-2抑制剂数据集上,模型实现了平均0.751的皮尔逊相关系数,证实其在该靶点上的可靠性。
- Gly143被确定为最有利于形成氢键的残基,其次为Cys145、Glu166和His163。
- 共发现45种抑制剂与催化性Cys145残基形成共价键,主要通过二碳单氧化物基团。
- 在排名前十的抑制剂中仅有两个(5rg1和6w63)为非共价抑制剂,表明对共价抑制具有强烈偏好。
- 片段分解揭示了来自顶级抑制剂的126种独特药效团片段,为从头药物设计提供了可操作的见解。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。