Skip to main content
QUICK REVIEW

[论文解读] How Evaluation Choices Distort the Outcome of Generative Drug Discovery

Rıza Özçelik, Francesca Grisoni|arXiv (Cornell University)|Dec 24, 2024
Biotechnology and Related FieldsMedicine被引用 3
一句话总结

本文批判性地评估了生成式药物发现中的评估实践,揭示了库大小和采样策略系统性地扭曲性能指标,导致对分子质量的错误结论。通过在10⁹个全新设计的大规模分析中,识别出关键陷阱,引入标准化评估工具('treasures'),并提出可操作的指导原则('ways out'),以改进模型基准测试和前瞻性分子选择。

ABSTRACT

"How to evaluate the de novo designs proposed by a generative model?" Despite the transformative potential of generative deep learning in drug discovery, this seemingly simple question has no clear answer. The absence of standardized guidelines challenges both the benchmarking of generative approaches and the selection of molecules for prospective studies. In this work, we take a fresh - critical and constructive - perspective on de novo design evaluation. By training chemical language models, we analyze approximately 1 billion molecule designs and discover principles consistent across different neural networks and datasets. We uncover a key confounder: the size of the generated molecular library significantly impacts evaluation outcomes, often leading to misleading model comparisons. We find increasing the number of designs as a remedy and propose new and compute-efficient metrics to compute at large-scale. We also identify critical pitfalls in commonly used metrics - such as uniqueness and distributional similarity - that can distort assessments of generative performance. To address these issues, we propose new and refined strategies for reliable model comparison and design evaluation. Furthermore, when examining molecule selection and sampling strategies, our findings reveal the constraints to diversify the generated libraries and draw new parallels and distinctions between deep learning and drug discovery. We anticipate our findings to help reshape evaluation pipelines in generative drug discovery, paving the way for more reliable and reproducible generative modeling approaches.

研究动机与目标

  • 识别并揭示当前生成式药物发现模型评估实践中存在的关键缺陷。
  • 研究库大小和采样策略如何扭曲对生成分子质量的感知。
  • 开发标准化、可靠的评估工具和指标,以客观比较生成模型。
  • 提供可操作的指导原则('ways out'),用于选择高质量、新颖且多样化的分子,以供前瞻性实验验证。
  • 通过建立系统化的评估框架,加强分子设计与深度学习之间的整合。

提出的方法

  • 在ChEMBLv33的150万个SMILES上训练三种最先进的化学语言模型(LSTM、GPT、S4),并在3个生物活性靶标(DRD3、PIN1、VDR)上使用5组随机划分的各320个活性分子进行微调。
  • 每种模型、采样策略和超参数设置生成100万个SMILES,总计约10⁹个全新设计分子,涵盖22种采样配置(温度、top-k、top-p)。
  • 使用全面的指标套件评估设计结果:语法正确性、唯一性、新颖性、FCD(Fréchet ChemNet距离)、FDD(Fréchet描述符距离)、子结构相似性(ECFP上的Tanimoto系数)、内部多样性(球体排斥法)以及设计似然度(对数概率乘积)。
  • 系统性地将库大小从100调整至1,000,000,以评估其对评估结果的影响,特别是对相似性和多样性指标的影响。
  • 采用温度、top-k和top-p采样策略,探究生成策略如何影响分子质量和分布保真度。
  • 在计算FDD前,对分子描述符(logP、MW、HBD、环数、PSA)进行最小-最大归一化,并使用RDKit进行指纹和描述符计算。

实验结果

研究问题

  • RQ1生成分子库的大小在多大程度上系统性地扭曲了全新药物发现中分子质量的评估?
  • RQ2采样超参数(温度、top-k、top-p)在多大程度上扭曲了对生成分子相似性和多样性的感知?
  • RQ3广泛使用的评估指标(如FCD、FDD、新颖性)与实际分子质量和生物相关性之间的相关性如何?
  • RQ4能否建立标准化、可复现的评估协议,以实现在不同生成模型之间的公平基准测试?
  • RQ5在生成式药物发现的评估领域中,关键的'陷阱'、'宝藏'和'出路'分别是什么?

主要发现

  • 库大小显著影响评估结果:较小的库(如1,000个设计)系统性地高估与训练集的相似性,低估多样性,从而导致对模型性能的错误信心。
  • Fréchet ChemNet距离(FCD)和Fréchet描述符距离(FDD)对库大小敏感,当比较1,000个与100,000个设计时,FCD值可下降高达30%,即使模型质量保持不变。
  • Top-1(贪婪)采样产生高度相似的分子,多样性极低(平均Tanimoto相似性 >0.95),而提高温度或采用top-p采样虽可提升多样性,但可能生成无效或低质量分子。
  • 新颖性和唯一性指标对库大小极为敏感:当库大小从1,000增至100,000时,即使对于高质量模型,新颖性比率也可能下降高达50%。
  • 设计似然度(对数概率)与生物相关性相关性较差,表明高概率分子不一定是前瞻性研究的更优候选。
  • 本研究识别出22种不同的采样配置,其评估结果存在显著差异,凸显了建立标准化协议以避免误导性基准测试的必要性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。