Skip to main content
QUICK REVIEW

[Paper Review] It Takes Two to Tango: Directly Optimizing for Constrained Synthesizability in Generative Molecular Design

Jeff Guo, Philippe Schwaller|arXiv (Cornell University)|Oct 15, 2024
Chemical Synthesis and AnalysisBiochemistry, Genetics and Molecular Biology3 citations
TL;DR

This paper introduces TANGO, a novel dense reward function that directly optimizes generative molecular models for constrained synthesizability using chemistry principles, enabling reinforcement learning to generate molecules that satisfy multi-parameter optimization goals while enforcing specific commercial building blocks. It is the first framework to achieve this, demonstrating success across starting-material, intermediate, and divergent synthesis constraints with state-of-the-art performance.

ABSTRACT

Constrained synthesizability is an unaddressed challenge in generative molecular design. In particular, designing molecules satisfying multi-parameter optimization objectives, while simultaneously being synthesizable and enforcing the presence of specific commercial building blocks in the synthesis. This is practically important for molecule re-purposing, sustainability, and efficiency. In this work, we propose a novel reward function called TANimoto Group Overlap (TANGO), which uses chemistry principles to transform a sparse reward function into a dense and learnable reward function -- crucial for reinforcement learning. TANGO can augment general-purpose molecular generative models to directly optimize for constrained synthesizability while simultaneously optimizing for other properties relevant to drug discovery using reinforcement learning. Our framework is general and addresses starting-material, intermediate, and divergent synthesis constraints. Contrary to most existing works in the field, we show that incentivizing a general-purpose (without any inductive biases) model is a productive approach to navigating challenging optimization scenarios. We demonstrate this by showing that the trained models explicitly learn a desirable distribution. Our framework is the first generative approach to tackle constrained synthesizability.

Motivation & Objective

  • To address the unmet challenge of directly optimizing generative molecular models for constrained synthesizability, including starting-material, intermediate, and divergent synthesis constraints.
  • To develop a dense, learnable reward function that transforms sparse rewards into actionable signals for reinforcement learning.
  • To enable general-purpose generative models to learn synthesizable molecular distributions without inductive biases, while optimizing for drug discovery-relevant properties.
  • To demonstrate that incentivizing a general-purpose model—rather than constraining its architecture—is a productive approach to solving complex multi-objective molecular design problems.
  • To show that the framework can generate molecules with high docking scores and high QED values, while ensuring enforceable building blocks are used in the synthetic routes.

Proposed method

  • Propose TANimoto Group Overlap (TANGO), a chemistry-inspired dense reward function that quantifies the overlap between functional groups in the generated molecule and those in the enforced building blocks.
  • Use retrosynthesis models as oracles to evaluate synthesizability and guide the reward function, ensuring generated molecules are synthetically feasible.
  • Integrate TANGO into a reinforcement learning framework that jointly optimizes for multiple molecular properties (e.g., docking score, QED) and synthesizability.
  • Apply the method to diverse synthesis constraints: starting-material, intermediate, and divergent synthesis, with enforced building blocks.
  • Use a general-purpose generative model (without inductive biases) and let it learn synthesizability through reward shaping, rather than architectural constraints.
  • Validate the approach using multiple configurations, including varying oracle budgets and QED enforcement, across 10 random seeds.
Figure 1: TANGO guides the generation of molecules directly optimized for constrained synthesizability with enforced building blocks while simultaneously optimizing other properties. Our method generalizes across starting-material, intermediate, and divergent synthesis constraints.
Figure 1: TANGO guides the generation of molecules directly optimized for constrained synthesizability with enforced building blocks while simultaneously optimizing other properties. Our method generalizes across starting-material, intermediate, and divergent synthesis constraints.

Experimental results

Research questions

  • RQ1Can a general-purpose generative model be effectively incentivized to learn synthesizable molecular distributions without architectural inductive biases?
  • RQ2Can a dense, chemistry-grounded reward function like TANGO effectively guide reinforcement learning toward molecules that satisfy multi-parameter optimization and enforce specific building blocks?
  • RQ3Is it possible to simultaneously optimize for drug discovery-relevant properties (e.g., docking score, QED) and constrained synthesizability in a single training regime?
  • RQ4How does the performance of the framework vary across different types of synthesis constraints—starting-material, intermediate, and divergent synthesis?
  • RQ5Does increasing the oracle budget improve the success rate of generating molecules with enforced building blocks, and is this scalable in practice?

Key findings

  • The TANGO framework successfully generated molecules with enforced building blocks across all tested configurations, including divergent synthesis, with 4 out of 10 seeds achieving at least one valid molecule under a 10,000 oracle budget.
  • With a 15,000 oracle budget, the success rate increased to 5 out of 10 seeds for divergent blocks, demonstrating that increased compute improves performance.
  • The average docking score of generated molecules with enforced blocks was -8.48 ± 0.25 (M=2694) under the 10,000 budget, indicating high binding affinity potential.
  • The average QED value of molecules in the highest docking score interval (DS < -10) was 0.84 ± 0.10, indicating high drug-likeness.
  • In the absence of QED enforcement, the success rate dropped to 3 out of 10 seeds, but increased to 4 out of 10 with a 15,000 budget, showing that QED is a useful but not essential constraint.
  • The average number of reaction steps for successfully generated molecules was 3.68 ± 1.08 under the 10,000 budget, indicating feasible synthetic routes.
Figure 2: TANGO reward function: the maximum similarity between every non-root node (generated molecule) molecule and the set of enforced building blocks. Every synthesizable generated molecule returns a non-zero reward.
Figure 2: TANGO reward function: the maximum similarity between every non-root node (generated molecule) molecule and the set of enforced building blocks. Every synthesizable generated molecule returns a non-zero reward.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.