Skip to main content
QUICK REVIEW

[论文解读] TopP-S: Persistent homology based multi-task deep neural networks for simultaneous predictions of partition coefficient and aqueous solubility

Kedi Wu, Zhixiong Zhao|arXiv (Cornell University)|Dec 9, 2017
Topological and Geometric Data Analysis被引用 4
一句话总结

本论文提出TopP-S,一种多任务深度学习框架,利用元素特异性持久同调(ESPH)同时预测小分子的水溶性和分配系数(logP)。通过将分子拓扑结构编码为可扩展的多尺度拓扑不变量,该方法实现了高度准确的预测,在基准数据集上通过联合学习相关理化性质,优于当前最先进模型。

ABSTRACT

Aqueous solubility and partition coefficient are important physical properties of small molecules. Accurate theoretical prediction of aqueous solubility and partition coefficient plays an important role in drug design and discovery. The prediction accuracy depends crucially on molecular descriptors which are typically derived from theoretical understanding of the chemistry and physics of small molecules. The present work introduces an algebraic topology based method, called element specific persistent homology (ESPH), as a new representation of small molecules that is entirely different from conventional chemical and/or physical representations. ESPH describes molecular properties in terms of multiscale and multicomponent topological invariants. Such topological representation is systematical, comprehensive, and scalable with respect to molecular size and composition variations. However, it cannot be literally translated into a physical interpretation. Fortunately, it is readily suitable for machine learning methods, rendering topological learning algorithms. Due to the inherent correlation between solubility and partition coefficient, a uniform ESPH representation is developed for both properties, which facilitates multi-task deep neural networks for their simultaneous predictions. This strategy leads to more accurate prediction of relatively small data sets. A total of six data sets is considered in the present work to validate the proposed topological and multi-task deep learning approaches. It is demonstrate that the proposed approaches achieve some of the most accurate predictions of aqueous solubility and partition coefficient. Our software is available online at {\url{http://weilab.math.msu.edu/TopP-S/}}

研究动机与目标

  • 开发一种基于代数拓扑的新分子表征方法,以捕捉多尺度、多组分的结构特征。
  • 通过数据高效、拓扑感知的机器学习方法,解决在药物设计中准确预测logP与水溶性——关键理化性质——的挑战。
  • 通过多任务学习联合建模logP与水溶性,利用二者固有的相关性。
  • 证明基于持久同调的拓扑描述符在预测精度上可超越传统分子描述符。

提出的方法

  • 采用元素特异性持久同调(ESPH)从分子几何结构中生成多尺度、多组分的拓扑不变量,编码原子身份与成键模式。
  • 构建元素特异性拓扑描述符(ESTDs)作为ESPH输出的紧凑、可解释性总结,保留化学与拓扑信息。
  • 在多任务学习框架下,将ESTDs与深度神经网络结合,联合预测logP与水溶性。
  • 应用集成方法(如梯度提升、随机森林)与深度神经网络,评估在多样化数据集上的预测性能。
  • 采用10折交叉验证与留一法验证,评估模型的泛化能力与鲁棒性。
  • 利用任务间的共享表征,通过相关性质带来的归纳偏差,提升在小样本或数据有限数据集上的性能。

实验结果

研究问题

  • RQ1元素特异性持久同调(ESPH)能否提供一种系统化、可扩展且信息丰富的分子表征,以捕捉预测理化性质所必需的化学与物理特征?
  • RQ2与单任务模型相比,联合学习logP与水溶性是否能显著提升预测精度?
  • RQ3拓扑描述符(ESTDs)在预测logP与水溶性方面,相较于传统2D分子描述符,优势程度如何?
  • RQ4尽管缺乏直接的物理解释,基于拓扑的表征是否仍能在基准数据集上实现最先进性能?

主要发现

  • 所提出的基于ESPH的拓扑表征,以ESTDs形式编码,已在六个基准数据集上实现了logP与水溶性预测的最先进精度。
  • 在共享ESTD表征上训练的多任务深度神经网络(MT-DNN)显著优于单任务模型,尤其在小样本数据集上,得益于共享归纳偏差带来的泛化性能提升。
  • 该方法在FDA与Star数据集上对logP预测实现了最先进性能,报告的RMSE值低于0.5 log单位。
  • 在传统2D描述符基础上加入ESTDs可进一步提升预测精度,表明信息互补性。
  • 该框架在多样化数据分布下表现出鲁棒性,交叉验证与测试集评估中均保持一致的性能提升。
  • 在线服务器 http://weilab.math.msu.edu/TopP-S/ 提供了模型的可访问接口,支持可复现性与在药物发现流程中的实际应用。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。