[Paper Review] It Takes Two to Tango: Directly Optimizing for Constrained Synthesizability in Generative Molecular Design
This paper introduces TANGO, a novel dense reward function that directly optimizes generative molecular models for constrained synthesizability using chemistry principles, enabling reinforcement learning to generate molecules that satisfy multi-parameter optimization goals while enforcing specific commercial building blocks. It is the first framework to achieve this, demonstrating success across starting-material, intermediate, and divergent synthesis constraints with state-of-the-art performance.
Constrained synthesizability is an unaddressed challenge in generative molecular design. In particular, designing molecules satisfying multi-parameter optimization objectives, while simultaneously being synthesizable and enforcing the presence of specific commercial building blocks in the synthesis. This is practically important for molecule re-purposing, sustainability, and efficiency. In this work, we propose a novel reward function called TANimoto Group Overlap (TANGO), which uses chemistry principles to transform a sparse reward function into a dense and learnable reward function -- crucial for reinforcement learning. TANGO can augment general-purpose molecular generative models to directly optimize for constrained synthesizability while simultaneously optimizing for other properties relevant to drug discovery using reinforcement learning. Our framework is general and addresses starting-material, intermediate, and divergent synthesis constraints. Contrary to most existing works in the field, we show that incentivizing a general-purpose (without any inductive biases) model is a productive approach to navigating challenging optimization scenarios. We demonstrate this by showing that the trained models explicitly learn a desirable distribution. Our framework is the first generative approach to tackle constrained synthesizability.
Motivation & Objective
- To address the unmet challenge of directly optimizing generative molecular models for constrained synthesizability, including starting-material, intermediate, and divergent synthesis constraints.
- To develop a dense, learnable reward function that transforms sparse rewards into actionable signals for reinforcement learning.
- To enable general-purpose generative models to learn synthesizable molecular distributions without inductive biases, while optimizing for drug discovery-relevant properties.
- To demonstrate that incentivizing a general-purpose model—rather than constraining its architecture—is a productive approach to solving complex multi-objective molecular design problems.
- To show that the framework can generate molecules with high docking scores and high QED values, while ensuring enforceable building blocks are used in the synthetic routes.
Proposed method
- Propose TANimoto Group Overlap (TANGO), a chemistry-inspired dense reward function that quantifies the overlap between functional groups in the generated molecule and those in the enforced building blocks.
- Use retrosynthesis models as oracles to evaluate synthesizability and guide the reward function, ensuring generated molecules are synthetically feasible.
- Integrate TANGO into a reinforcement learning framework that jointly optimizes for multiple molecular properties (e.g., docking score, QED) and synthesizability.
- Apply the method to diverse synthesis constraints: starting-material, intermediate, and divergent synthesis, with enforced building blocks.
- Use a general-purpose generative model (without inductive biases) and let it learn synthesizability through reward shaping, rather than architectural constraints.
- Validate the approach using multiple configurations, including varying oracle budgets and QED enforcement, across 10 random seeds.

Experimental results
Research questions
- RQ1Can a general-purpose generative model be effectively incentivized to learn synthesizable molecular distributions without architectural inductive biases?
- RQ2Can a dense, chemistry-grounded reward function like TANGO effectively guide reinforcement learning toward molecules that satisfy multi-parameter optimization and enforce specific building blocks?
- RQ3Is it possible to simultaneously optimize for drug discovery-relevant properties (e.g., docking score, QED) and constrained synthesizability in a single training regime?
- RQ4How does the performance of the framework vary across different types of synthesis constraints—starting-material, intermediate, and divergent synthesis?
- RQ5Does increasing the oracle budget improve the success rate of generating molecules with enforced building blocks, and is this scalable in practice?
Key findings
- The TANGO framework successfully generated molecules with enforced building blocks across all tested configurations, including divergent synthesis, with 4 out of 10 seeds achieving at least one valid molecule under a 10,000 oracle budget.
- With a 15,000 oracle budget, the success rate increased to 5 out of 10 seeds for divergent blocks, demonstrating that increased compute improves performance.
- The average docking score of generated molecules with enforced blocks was -8.48 ± 0.25 (M=2694) under the 10,000 budget, indicating high binding affinity potential.
- The average QED value of molecules in the highest docking score interval (DS < -10) was 0.84 ± 0.10, indicating high drug-likeness.
- In the absence of QED enforcement, the success rate dropped to 3 out of 10 seeds, but increased to 4 out of 10 with a 15,000 budget, showing that QED is a useful but not essential constraint.
- The average number of reaction steps for successfully generated molecules was 3.68 ± 1.08 under the 10,000 budget, indicating feasible synthetic routes.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.