[Paper Review] How Evaluation Choices Distort the Outcome of Generative Drug Discovery
This paper critically evaluates evaluation practices in generative drug discovery, revealing that library size and sampling strategies systematically bias performance metrics, leading to false conclusions about molecular quality. Using large-scale analysis across 10⁹ de novo designs, it identifies key pitfalls, introduces standardized evaluation tools ('treasures'), and proposes actionable guidelines ('ways out') to improve model benchmarking and prospective molecule selection.
"How to evaluate the de novo designs proposed by a generative model?" Despite the transformative potential of generative deep learning in drug discovery, this seemingly simple question has no clear answer. The absence of standardized guidelines challenges both the benchmarking of generative approaches and the selection of molecules for prospective studies. In this work, we take a fresh - critical and constructive - perspective on de novo design evaluation. By training chemical language models, we analyze approximately 1 billion molecule designs and discover principles consistent across different neural networks and datasets. We uncover a key confounder: the size of the generated molecular library significantly impacts evaluation outcomes, often leading to misleading model comparisons. We find increasing the number of designs as a remedy and propose new and compute-efficient metrics to compute at large-scale. We also identify critical pitfalls in commonly used metrics - such as uniqueness and distributional similarity - that can distort assessments of generative performance. To address these issues, we propose new and refined strategies for reliable model comparison and design evaluation. Furthermore, when examining molecule selection and sampling strategies, our findings reveal the constraints to diversify the generated libraries and draw new parallels and distinctions between deep learning and drug discovery. We anticipate our findings to help reshape evaluation pipelines in generative drug discovery, paving the way for more reliable and reproducible generative modeling approaches.
Motivation & Objective
- To identify and expose critical flaws in current evaluation practices for generative drug discovery models.
- To investigate how library size and sampling strategies distort the perceived quality of generated molecules.
- To develop standardized, reliable evaluation tools and metrics for comparing generative models objectively.
- To provide actionable guidelines ('ways out') for selecting high-quality, novel, and diverse molecules for prospective experimental validation.
- To strengthen the integration between molecular design and deep learning by establishing a systematic evaluation framework.
Proposed method
- Trained three state-of-the-art chemical language models (LSTM, GPT, S4) on 1.5M SMILES from ChEMBLv33 and fine-tuned them on 3 bioactive targets (DRD3, PIN1, VDR) using 5 random splits of 320 active molecules each.
- Generated 1 million SMILES per model, sampling strategy, and hyperparameter setting, totaling ~10⁹ de novo molecules across 22 sampling configurations (temperature, top-k, top-p).
- Evaluated designs using a comprehensive suite of metrics: syntactic validity, uniqueness, novelty, FCD (Fréchet ChemNet Distance), FDD (Fréchet Descriptor Distance), substructure similarity (Tanimoto on ECFPs), internal diversity (sphere exclusion), and design likelihood (log-probability product).
- Systematically varied library sizes from 100 to 1,000,000 to assess the impact on evaluation outcomes, particularly on similarity and diversity metrics.
- Used temperature, top-k, and top-p sampling strategies to explore how generation strategy affects molecular quality and distributional fidelity.
- Applied min-max normalization to molecular descriptors (logP, MW, HBD, rings, PSA) before computing FDD, and used RDKit for fingerprint and descriptor computation.
Experimental results
Research questions
- RQ1How does the size of the generated molecule library systematically bias the evaluation of molecular quality in de novo drug discovery?
- RQ2To what extent do sampling hyperparameters (temperature, top-k, top-p) distort the perceived similarity and diversity of generated molecules?
- RQ3How do widely used evaluation metrics (e.g., FCD, FDD, novelty) correlate with actual molecular quality and biological relevance?
- RQ4Can standardized, reproducible evaluation protocols be established to enable fair benchmarking across generative models?
- RQ5What are the key 'traps', 'treasures', and 'ways out' in the evaluation landscape of generative drug discovery?
Key findings
- Library size significantly biases evaluation: smaller libraries (e.g., 1,000 designs) systematically overestimate similarity to the training set and underestimate diversity, leading to false confidence in model performance.
- The Fréchet ChemNet Distance (FCD) and Fréchet Descriptor Distance (FDD) are sensitive to library size, with FCD values decreasing by up to 30% when comparing 1,000 to 100,000 designs, even when model quality is unchanged.
- Top-1 (greedy) sampling produces highly similar molecules with low diversity (mean Tanimoto similarity >0.95), while higher temperature or top-p sampling increases diversity but risks generating invalid or low-quality molecules.
- Novelty and uniqueness metrics are highly sensitive to library size: novelty rates drop by up to 50% when increasing library size from 1,000 to 100,000, even for high-quality models.
- Design likelihood (log-probability) correlates poorly with biological relevance, indicating that high-probability molecules are not necessarily better candidates for prospective studies.
- The study identifies 22 distinct sampling configurations that yield significantly different evaluation outcomes, underscoring the need for standardized protocols to avoid misleading benchmarking.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.