[论文解读] Beyond the Training Domain: Robust Generative Transition State Models for Unseen Chemistry
该论文在新的元素和催化化学中对生成型过渡态(TS)模型进行基准测试,揭示泛化极限,并引入基于构象体(conformer)的自监督预训练,以提升对未见化学体系的TS预测,从而降低微调所需数据量。
Transition states (TSs) govern the rates and outcomes of chemical reactions, making their accurate prediction a central challenge in computational chemistry. Although recent machine-learning models achieve near chemical accuracy in the prediction of TS structures and the associated reaction barriers for small organic reactions, their ability to generalize beyond the training domain remains largely unexplored. Here, we introduce targeted benchmarks to probe chemical and structural novelty in generative TS prediction. Building on Transition1x, a large-scale dataset of reactions involving small organic molecules, we construct curated extensions incorporating controlled elemental substitutions and diverse transition-metal complexes (TMC). These benchmarks reveal fundamental limitations of generative models in the generalization to previously unseen elements. As a result, they produce unphysical geometries and large energetic errors, even for reactions structurally similar to well-predicted organic systems. To address this challenge, we introduce a self-supervised pretraining strategy based on equilibrium conformers that exposes generative TS models to novel chemical environments prior to targeted fine-tuning. Across the newly proposed benchmarks, self-supervised pretraining substantially improves TS prediction for previously unseen systems, lowering the median RMSD of TS geometries on T1x-TMC reactions from 0.39 to 0.19 $\mathring{A}$ and reducing fine-tuning data requirements by up to 75%, enabling reliable performance even in low-data regimes. Overall, the integration of generative TS models with self-supervised pseudo-reaction pretraining provides an efficient, scalable, and chemically robust framework for elucidating TSs well beyond the small organic molecule domain, establishing a foundation for investigating complex and catalytically relevant reaction landscapes.
研究动机与目标
- 评估最先进的生成型TS模型在小分子以外的普遍性。
- 开发引入元素新颖性和过渡金属络合物(TMC)化学的基准。
- 评估现有模型在分布外化学中的局限性和故障模式。
- 提出基于平衡构象的自监督预训练策略,以提升迁移能力和数据效率。
提出的方法
- 通过用同族元素替换同一周期内的单个原子并在P-RFO+IRC重新优化TS,创建Transition1x-2p3p4p。
- 将Transition1x TS嵌入到十个催化相关的过渡金属络合物中,并在GFN2-xTB中优化,创建Transition1x-TMC。
- 在新基准上评估基线模型(React-OT和AEFM),以识别新元素引入时的性能下降。
- 通过从平衡构象构建伪反应(最高能量为TS,中间为反应物,最低为产物)来应用基于构象体的自监督预训练。
- 在目标数据集上微调预训练模型,并评估TS几何精度(RMSD)和能量误差的提升。
- 演示通过选择性重新优化实现对DFT级别的迁移,并比较GFN2-xTB与DFT能量之间的差异。

实验结果
研究问题
- RQ1现有的生成型TS模型在遇到未见元素和新反应机理的反应时表现如何?
- RQ2在将TS预测外推到分布外的化学体系时,主要的故障模式是什么?
- RQ3基于构象体的自监督预训练能否提高对未见化学体系中TS预测的泛化性和数据效率?
- RQ4在多大程度上可以将半经验(GFN2-xTB)与DFT级数据结合,以在保持准确性的同时实现高通量探索?
主要发现
- 当引入新元素时,生成型TS模型(React-OT、AEFM)的性能快速下降(Transition1x-TMC中最多可有两个新元素)。
- 对于Transition1x-2p3p4p,未加权RMSD从0.04 Å(HCNO)增至引入一个新元素时的0.18 Å;对于Transition1x-TMC,中位RMSD升至0.39 Å(对比HCNO的0.05 Å)。
- 在平衡构象上的自监督预训练显著提升TS预测,在数据集不同情况下中位RMSD降至0.10–0.19 Å,并将微调数据需求减少最多75%。
- 使用伪反应进行预训练实现数据高效迁移,在仅使用真实反应的一小部分数据下接近完全训练的性能(如25–50%数据)。
- 以GFN2-xTB为可扩展基础的混合方法与DFT级TS能量具有合理的一致性(ΔE_TS在各数据集内约25%范围内),选择的预测可以显著提高收敛到DFT TS结构的概率。
- DFT级构象预训练进一步提升精度(例如将Transition1x-TMC的RMSD从0.47降至0.42 Å,使用1500个伪反应)。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。