[论文解读] Learning Together: Towards foundational models for machine learning interatomic potentials with meta-learning
本文提出使用元学习训练基础机器学习原子间势(MLIPs),使其能够联合学习多个量子力学(QM)数据集,这些数据集具有不同理论水平。通过实现对新分子的快速适应,元学习提升了泛化能力,降低了误差,并增强了势能面的平滑性,在药物样分子如3BPA上的表现优于标准迁移学习。
The development of machine learning models has led to an abundance of datasets containing quantum mechanical (QM) calculations for molecular and material systems. However, traditional training methods for machine learning models are unable to leverage the plethora of data available as they require that each dataset be generated using the same QM method. Taking machine learning interatomic potentials (MLIPs) as an example, we show that meta-learning techniques, a recent advancement from the machine learning community, can be used to fit multiple levels of QM theory in the same training process. Meta-learning changes the training procedure to learn a representation that can be easily re-trained to new tasks with small amounts of data. We then demonstrate that meta-learning enables simultaneously training to multiple large organic molecule datasets. As a proof of concept, we examine the performance of a MLIP refit to a small drug-like molecule and show that pre-training potentials to multiple levels of theory with meta-learning improves performance. This difference in performance can be seen both in the reduced error and in the improved smoothness of the potential energy surface produced. We therefore show that meta-learning can utilize existing datasets with inconsistent QM levels of theory to produce models that are better at specializing to new datasets. This opens new routes for creating pre-trained, foundational models for interatomic potentials.
研究动机与目标
- 解决整合大规模、异构QM数据集(理论水平不一致)以训练MLIPs的挑战。
- 实现可迁移、通用的原子间势,可快速适应新分子体系。
- 证明元学习在整合具有不同QM理论水平的多个数据集时,可优于标准迁移学习。
- 建立利用现有多保真度数据集预训练基础模型的路径,用于原子间势。
- 强调标准化数据格式的必要性,以实现元学习在更广阔化学空间的扩展。
提出的方法
- 使用元学习同时在多个具有不同理论水平(如DFT、CCSD(T))的QM数据集上训练单一模型。
- 采用双层优化框架,使模型学习一个可在各类任务间泛化的初始化,从而实现少量样本下的快速适应。
- 在五个大型有机分子数据集(QM7-x、QMugs、ANI-1x、Transition-1x和GEOM)上进行预训练。
- 对不同QM理论水平的数据集实施线性缩放,以对齐能量,确保一致的训练信号。
- 在小规模目标数据集(如3BPA)上微调元学习得到的模型,以评估其专业化性能。
- 采用消息传递神经网络架构(如SchNet或PhysNet)作为MLIP的基础模型。

实验结果
研究问题
- RQ1元学习能否实现对具有不一致理论水平的多个QM数据集的MLIP联合训练?
- RQ2通过元学习在多样化数据集上预训练,是否能提升对未见分子体系的泛化能力与准确性?
- RQ3与标准迁移学习相比,基于元学习的适应在势能面的误差与平滑性方面表现如何?
- RQ4元学习模型能否在极少微调下实现对3BPA等小而复杂分子的更优性能?
- RQ5整合具有不同QM理论水平的数据集存在哪些实际限制?在何种情况下额外数据具有益处?
主要发现
- 在QM7-x、QMugs、ANI-1x、Transition-1x和GEOM等多个数据集上预训练的元学习模型,在微调3BPA分子时表现优于标准迁移学习。
- 元学习模型生成的势能面显著比标准迁移学习更平滑,减少了噪声并提升了物理一致性。
- 元学习模型在3BPA测试集上实现了更低的平均绝对误差(MAE),证明其具备更高的精度与泛化能力。
- 尽管在多个数据集上预训练提升了对多样化体系的性能,但直接从ANI-1x微调至CCSD(T)时误差最低,表明特定任务中数据一致性至关重要。
- 元学习可有效利用现有异构数据集,无需进行新的QM计算,从而加速可迁移MLIP的开发。
- 本研究强调了标准化数据格式的紧迫需求,以推动元学习在更广泛材料与分子科学领域的扩展。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。