Skip to main content
QUICK REVIEW

[Paper Review] Deep learning for molecular generation and optimization - a review of the state of the art

Daniel C. Elton, Zois Boukouvalas|arXiv (Cornell University)|Mar 11, 2019
Machine Learning in Materials ScienceMaterials Science66 references20 citations
TL;DR

This review synthesizes recent advances in deep generative modeling for molecular generation and optimization, evaluating four key approaches—recursive neural networks, autoencoders, GANs, and reinforcement learning. It highlights the shift toward graph and 3D molecular representations, the critical role of reward function design, and the superiority of adversarial and reinforcement learning over maximum likelihood training for generating drug-like molecules.

ABSTRACT

In the space of only a few years, deep generative modeling has revolutionized how we think of artificial creativity, yielding autonomous systems which produce original images, music, and text. Inspired by these successes, researchers are now applying deep generative modeling techniques to the generation and optimization of molecules - in our review we found 45 papers on the subject published in the past two years. These works point to a future where such systems will be used to generate lead molecules, greatly reducing resources spent downstream synthesizing and characterizing bad leads in the lab. In this review we survey the increasingly complex landscape of models and representation schemes that have been proposed. The four classes of techniques we describe are recursive neural networks, autoencoders, generative adversarial networks, and reinforcement learning. After first discussing some of the mathematical fundamentals of each technique, we draw high level connections and comparisons with other techniques and expose the pros and cons of each. Several important high level themes emerge as a result of this work, including the shift away from the SMILES string representation of molecules towards more sophisticated representations such as graph grammars and 3D representations, the importance of reward function design, the need for better standards for benchmarking and testing, and the benefits of adversarial training and reinforcement learning over maximum likelihood based training.

Motivation & Objective

  • To survey the state of the art in deep generative modeling for molecular generation and optimization.
  • To analyze the strengths and limitations of four major deep learning techniques: recursive neural networks, autoencoders, GANs, and reinforcement learning.
  • To identify emerging trends such as the move from SMILES strings to graph and 3D representations.
  • To emphasize the importance of reward function design and the need for standardized benchmarking in molecular generation research.

Proposed method

  • The paper conducts a comprehensive review of 45 recent papers (2021–2023) on deep generative modeling for molecular generation.
  • It categorizes and compares four main deep learning techniques: recursive neural networks, autoencoders, generative adversarial networks (GANs), and reinforcement learning.
  • It evaluates each method based on mathematical foundations, representation schemes (e.g., SMILES, graph grammars, 3D structures), and training objectives.
  • It contrasts maximum likelihood-based training with adversarial and reinforcement learning approaches, highlighting differences in optimization goals and outcome quality.
  • It discusses the role of reward functions in guiding molecular optimization toward desired chemical and biological properties.
  • It identifies key challenges such as lack of standard benchmarks and the need for better evaluation protocols in molecular generation.

Experimental results

Research questions

  • RQ1How do different deep generative models compare in their ability to generate novel, drug-like molecules?
  • RQ2What are the advantages and limitations of using SMILES strings versus graph or 3D representations in molecular generation?
  • RQ3How does reward function design influence the quality and novelty of generated molecules?
  • RQ4Why do adversarial and reinforcement learning methods outperform maximum likelihood-based training in molecular generation?
  • RQ5What are the current gaps in benchmarking and evaluation standards for molecular generation models?

Key findings

  • There has been a significant shift from SMILES string representations toward more sophisticated representations such as graph grammars and 3D molecular structures.
  • Adversarial training and reinforcement learning have shown superior performance compared to maximum likelihood-based training in generating high-quality, diverse, and property-optimized molecules.
  • Reward function design is a critical factor in guiding the generation of molecules with desired chemical and biological properties.
  • Despite rapid progress, a lack of standardized benchmarks and evaluation protocols remains a major obstacle to reliable comparison across models.
  • The field is advancing toward systems capable of autonomously generating lead compounds, reducing costly experimental screening in drug discovery.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.