Skip to main content
QUICK REVIEW

[论文解读] ToxTree: descriptor-based machine learning models for both hERG and Nav1.5 cardiotoxicity liability predictions

Issar Arab, Khaled Barakat|arXiv (Cornell University)|Dec 27, 2021
Computational Drug Discovery Methods被引用 4
一句话总结

ToxTree 引入了两种基于二维描述符的稳健机器学习模型——ToxTree-hERG 和 ToxTree-Nav1.5——通过 hERG 和 Nav1.5 通道阻断来预测心脏毒性风险。这两种模型分别在 8,380 和 1,550 种化合物的精选数据集上进行训练,达到最先进的性能表现,其中 Nav1.5 模型在外部测试集上的 Q4 为 74.9%,Q2 为 86.7%。

ABSTRACT

Drug-mediated blockade of the voltage-gated potassium channel(hERG) and the voltage-gated sodium channel (Nav1.5) can lead to severe cardiovascular complications. This rising concern has been reflected in the drug development arena, as the frequent emergence of cardiotoxicity from many approved drugs led to either discontinuing their use or, in some cases, their withdrawal from the market. Predicting potential hERG and Nav1.5 blockers at the outset of the drug discovery process can resolve this problem and can, therefore, decrease the time and expensive cost of developing safe drugs. One fast and cost-effective approach is to use in silico predictive methods to weed out potential hERG and Nav1.5 blockers at the early stages of drug development. Here, we introduce two robust 2D descriptor-based QSAR predictive models for both hERG and Nav1.5 liability predictions. The machine learning models were trained for both regression, predicting the potency value of a drug, and multiclass classification at three different potency cut-offs (i.e. 1$μ$M, 10$μ$M, and 30$μ$M), where ToxTree-hERG Classifier, a pipeline of Random Forest models, was trained on a large curated dataset of 8380 unique molecular compounds. Whereas ToxTree-Nav1.5 Classifier, a pipeline of kernelized SVM models, was trained on a large manually curated set of 1550 unique compounds retrieved from both ChEMBL and PubChem publicly available bioactivity databases. The proposed hERG inducer outperformed most metrics of the state-of-the-art published model and other existing tools. Additionally, we are introducing the first Nav1.5 liability predictive model achieving a Q4 = 74.9% and a binary classification of Q2 = 86.7% with MCC = 71.2% evaluated on an external test set of 173 unique compounds. The curated datasets used in this project are made publicly available to the research community.

研究动机与目标

  • 为解决因 hERG 和 Nav1.5 介导的心脏毒性导致药物研发中高失败率的问题。
  • 开发快速、低成本的体外计算模型,用于早期预测心脏毒性风险。
  • 创建公开可用的高质量精选数据集,涵盖 hERG 和 Nav1.5 生物活性数据。
  • 通过先进的机器学习流程,提升现有工具的预测准确性。

提出的方法

  • ToxTree-hERG 使用基于随机森林的流程,在来自公共生物活性数据库的 8,380 种独特化合物上进行训练。
  • ToxTree-Nav1.5 采用基于核函数的 SVM 流程,在来自 ChEMBL 和 PubChem 的 1,550 种人工筛选的 Nav1.5 化合物上进行训练。
  • 两种模型均使用二维分子描述符作为输入特征,以预测活性强度和分类结果。
  • 通过标准指标(包括 Q2、Q4 和 Matthews 相关系数(MCC))在外部测试集上对模型进行评估。
  • 在三个活性阈值(1 µM、10 µM 和 30 µM)下分别执行回归和多分类分类任务。
  • 用于训练的数据集已公开,以支持可复现性及后续研究。

实验结果

研究问题

  • RQ1基于描述符的机器学习模型能否在早期药物发现中准确预测 hERG 和 Nav1.5 心脏毒性风险?
  • RQ2ToxTree-hERG 和 ToxTree-Nav1.5 模型与现有最先进工具相比,在预测性能上表现如何?
  • RQ3使用精选的、公开可用的 Nav1.5 和 hERG 生物活性数据集,能够达到何种精度水平?
  • RQ4是否可以使用单一统一的流程,通过二维分子描述符有效建模 hERG 和 Nav1.5 的毒性风险?
  • RQ5Nav1.5 模型在外部测试集上使用多分类和二分类指标的性能如何?

主要发现

  • ToxTree-hERG 模型在预测 hERG 毒性风险方面优于大多数现有最先进模型。
  • ToxTree-Nav1.5 模型在 173 种独特化合物的外部测试集上达到 Q4 为 74.9%。
  • Nav1.5 模型在二分类任务中的 Q2 达到 86.7%,Matthews 相关系数(MCC)为 71.2%。
  • 本研究使用的精选数据集已公开,以支持社区使用和未来模型开发。
  • 模型在外部测试集上表现出强大的泛化能力,表明其在早期药物安全筛选中具有稳健性。
  • 将二维描述符与集成学习及核函数学习方法相结合,可实现对心脏毒性的高精度预测。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。