[Paper Review] Tartarus: A Benchmarking Platform for Realistic And Practical Inverse Molecular Design
TARTARUS provides a modular, realistic benchmark suite for inverse molecular design across materials, drugs, and reactions, comparing multiple generative models under practical constraints.
The efficient exploration of chemical space to design molecules with intended properties enables the accelerated discovery of drugs, materials, and catalysts, and is one of the most important outstanding challenges in chemistry. Encouraged by the recent surge in computer power and artificial intelligence development, many algorithms have been developed to tackle this problem. However, despite the emergence of many new approaches in recent years, comparatively little progress has been made in developing realistic benchmarks that reflect the complexity of molecular design for real-world applications. In this work, we develop a set of practical benchmark tasks relying on physical simulation of molecular systems mimicking real-life molecular design problems for materials, drugs, and chemical reactions. Additionally, we demonstrate the utility and ease of use of our new benchmark set by demonstrating how to compare the performance of several well-established families of algorithms. Surprisingly, we find that model performance can strongly depend on the benchmark domain. We believe that our benchmark suite will help move the field towards more realistic molecular design benchmarks, and move the development of inverse molecular design algorithms closer to designing molecules that solve existing problems in both academia and industry alike.
Motivation & Objective
- Motivate inverse molecular design and its real-world importance for drugs, catalysts, and materials.
- Provide a realistic, modular benchmark suite that reflects practical design problems using physical simulation workflows.
- Enable fair comparison of diverse algorithm families on multiple design domains.
- Offer detailed guidelines for dataset usage, evaluation, and reproducibility to accelerate adoption in chemistry and ML communities.
Proposed method
- Define four molecular design benchmark domains inspired by materials, drugs, and chemical reactions with curated reference datasets.
- Use physics-based and quantum-chemistry workflows (force fields, semi-empirical quantum chemistry, DFT) to compute target properties.
- Evaluate a wide range of generative models (REINVENT, SMILES-VAE, SELFIES-VAE, MoFlow, SMILES-LSTM-HC, SELFIES-LSTM-HC, GB-GA, JANUS) on each task.
- Adopt a constrained, resource-aware evaluation protocol (up to 5,000 proposals, 24 h runtime, 80/20 train/test split, five repeats).
- Provide installation/setup instructions and Supporting Information for reproducibility and extension.
Experimental results
Research questions
- RQ1How do different generative models perform across realistic molecular design benchmarks spanning materials, drugs, and reactions?
- RQ2To what extent does benchmark domain influence model performance and generalizability?
- RQ3Do simple genetic algorithms offer robust performance across domains compared to deep generative models?
- RQ4How do molecular representations (SMILES vs SELFIES) affect design outcomes and synthesizability?
Key findings
- Benchmark domain strongly affects model performance; no single model dominates across all tasks.
- GB-GA excels on the organic photovoltaics donor task while SMILES-LSTM-HC performs best on the acceptor task.
- In organic emitters, JANUS, GB-GA, and SELFIES-VAE generate competitive or superior results; some improvements in oscillator strengths observed.
- Protein-ligand benchmarks show no model consistently achieving top docking scores across all targets; SELFIES variants improve docking scores and SR in several cases.
- Reaction substrate benchmarks favor JANUS and GB-GA for outperforming best dataset molecules, with VAE models lagging on several objectives.
- VAEs generally exhibit slower training times and do not consistently outperform simpler or GA-based methods; MoFlow and REINVENT often demonstrate faster training and sampling.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.