[Paper Review] Symmetry-adapted generation of 3d point sets for the targeted discovery of molecules
The authors introduce G-SchNet, an autoregressive network that generates rotationally invariant 3D point sets (atoms with positions) for molecules, capturing 3D geometry and enabling bias toward properties like a small HOMO-LUMO gap. They validate on QM9, show proximity to equilibrium structures, and create novel molecules datasets.
Deep learning has proven to yield fast and accurate predictions of quantum-chemical properties to accelerate the discovery of novel molecules and materials. As an exhaustive exploration of the vast chemical space is still infeasible, we require generative models that guide our search towards systems with desired properties. While graph-based models have previously been proposed, they are restricted by a lack of spatial information such that they are unable to recognize spatial isomerism and non-bonded interactions. Here, we introduce a generative neural network for 3d point sets that respects the rotational invariance of the targeted structures. We apply it to the generation of molecules and demonstrate its ability to approximate the distribution of equilibrium structures using spatial metrics as well as established measures from chemoinformatics. As our model is able to capture the complex relationship between 3d geometry and electronic properties, we bias the distribution of the generator towards molecules with a small HOMO-LUMO gap - an important property for the design of organic solar cells.
Motivation & Objective
- Motivate geometry-aware molecule generation beyond graph-based methods to capture spatial isomerism and non-bonded interactions.
- Propose G-SchNet to generate 3D atomic positions and types with rotational and translational invariance.
- Demonstrate generation of novel, equilibrium-like molecules from QM9 and assess structural and spatial fidelity.
- Show how the generator can be biased toward desired electronic properties such as small HOMO-LUMO gaps.
- Provide datasets of novel generated molecules for further analysis and benchmarking.
Proposed method
- Autoregressive factorization of point-set distributions that is symmetry-adapted to rotations, translations, and local symmetries.
- Generation of next atom type and position using a distance-based probability conditioned on previously placed points.
- Utilization of auxiliary tokens (focus point and origin) to localize sampling and encode global geometry.
- SchNet-based neural network with continuous-filter convolutional layers to obtain rotation/translation-invariant atom features.
- Prediction of type distribution via a product of per-previous-point likelihoods (Eq. 4) and distance distribution via discretized bins (Eq. 3).
- Training via cross-entropy losses on type and distance distributions, with a stop token to end generation.
Experimental results
Research questions
- RQ1Can G-SchNet generate 3D molecular structures that resemble equilibrium geometries and reproduce structural statistics of QM9?
- RQ2Do generated structures exhibit correct spatial distributions (radial/ angular) compared to training data?
- RQ3Can the model be biased to increase molecules with desired electronic properties, e.g., small HOMO-LUMO gaps?
- RQ4How does G-SchNet compare to graph-based molecule generators in terms of validity, novelty, and structural characteristics?
- RQ5What datasets of novel generated molecules can be produced for further validation and benchmarking?
Key findings
- About 77% of generated molecules are valid after generation and valency checks.
- Generated molecules show RMSD medians around 0.21 Å for unseen/test data when compared to relaxed equilibrium structures.
- Radial distribution functions and angular distributions of generated molecules align well with QM9 training data, indicating faithful spatial statistics.
- Fine-tuning biased toward small HOMO-LUMO gaps increases the share of qualifying molecules from 7% to 43%.
- The authors introduce datasets with thousands of novel molecules not present in QM9 (over 9k new structures; >3.6k biased structures).
- Generated structures preserve atom/bond counts and resemble training data ring statistics when trained on filtered subsets (e.g., avoiding small rings).
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.