[Paper Review] Beyond the Training Domain: Robust Generative Transition State Models for Unseen Chemistry
The paper benchmarks generative transition-state (TS) models on new elemental and catalytic chemistry, reveals generalization limits, and introduces conformer-based self-supervised pretraining to improve TS predictions for unseen chemistry, reducing fine-tuning data needs.
Transition states (TSs) govern the rates and outcomes of chemical reactions, making their accurate prediction a central challenge in computational chemistry. Although recent machine-learning models achieve near chemical accuracy in the prediction of TS structures and the associated reaction barriers for small organic reactions, their ability to generalize beyond the training domain remains largely unexplored. Here, we introduce targeted benchmarks to probe chemical and structural novelty in generative TS prediction. Building on Transition1x, a large-scale dataset of reactions involving small organic molecules, we construct curated extensions incorporating controlled elemental substitutions and diverse transition-metal complexes (TMC). These benchmarks reveal fundamental limitations of generative models in the generalization to previously unseen elements. As a result, they produce unphysical geometries and large energetic errors, even for reactions structurally similar to well-predicted organic systems. To address this challenge, we introduce a self-supervised pretraining strategy based on equilibrium conformers that exposes generative TS models to novel chemical environments prior to targeted fine-tuning. Across the newly proposed benchmarks, self-supervised pretraining substantially improves TS prediction for previously unseen systems, lowering the median RMSD of TS geometries on T1x-TMC reactions from 0.39 to 0.19 $\mathring{A}$ and reducing fine-tuning data requirements by up to 75%, enabling reliable performance even in low-data regimes. Overall, the integration of generative TS models with self-supervised pseudo-reaction pretraining provides an efficient, scalable, and chemically robust framework for elucidating TSs well beyond the small organic molecule domain, establishing a foundation for investigating complex and catalytically relevant reaction landscapes.
Motivation & Objective
- Assess generalization of state-of-the-art generative TS models beyond small organic molecules.
- Develop benchmarks that introduce elemental novelty and transition-metal complex (TMC) chemistry.
- Evaluate limitations and failure modes of existing models on out-of-distribution chemistry.
- Propose a self-supervised pretraining strategy using equilibrium conformers to improve transferability and data efficiency.
Proposed method
- Create Transition1x-2p3p4p by substituting single atoms with same-group elements up to the third period and reoptimizing TS via P-RFO with IRC.
- Create Transition1x-TMC by embedding Transition1x TSs into ten catalytically relevant transition-metal complexes and optimizing at GFN2-xTB.
- Evaluate baseline models (React-OT and AEFM) on new benchmarks to identify performance degradation with novel elements.
- Apply conformer-based self-supervised pretraining by constructing pseudo-reactions from equilibrium conformers (highest-energy as TS, middle as reactant, lowest as product).
- Fine-tune pretrained models on target datasets and assess gains in TS geometry accuracy (RMSD) and energy errors.
- Demonstrate transferability to DFT-level via selective re-optimization and compare GFN2-xTB vs DFT energies.

Experimental results
Research questions
- RQ1How do existing generative TS models perform on reactions with unseen elements and new reaction mechanisms?
- RQ2What are the main failure modes when extrapolating TS predictions to out-of-distribution chemistry?
- RQ3Can conformer-based self-supervised pretraining improve generalization and data efficiency for TS prediction in unseen chemistry?
- RQ4To what extent can semi-empirical (GFN2-xTB) and DFT-level data be integrated to preserve accuracy while enabling high-throughput exploration?
Key findings
- Generative TS models (React-OT, AEFM) show rapid performance degradation as novel element types are introduced (up to two new elements in Transition1x-TMC).
- For Transition1x-2p3p4p, vanilla RMSD increases from 0.04 Å (HCNO) to 0.18 Å with one novel element; for Transition1x-TMC, median RMSD rises to 0.39 Å (vs 0.05 Å HCNO).
- Self-supervised pretraining on equilibrium conformers substantially improves TS predictions, reducing median RMSD to 0.10–0.19 Å depending on dataset, and decreases fine-tuning data needs by up to 75%.
- Pretraining with pseudo-reactions enables data-efficient transfer, achieving near-fully-trained performance with only a fraction of real reactions (e.g., 25–50% data).
- Hybrid approach using GFN2-xTB as a scalable foundation yields reasonable agreement with DFT-level TS energies (ΔE_TS within ~25% across datasets); selected predictions can converge to DFT TS structures with substantial success rates.
- DFT-level conformer pretraining further improves accuracy (e.g., reducing Transition1x-TMC RMSD from 0.47 to 0.42 Å with 1500 pseudo-reactions).

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.