Skip to main content
QUICK REVIEW

[論文レビュー] Learning Together: Towards foundational models for machine learning interatomic potentials with meta-learning

Alice E. A. Allen, Nicholas Lubbers|arXiv (Cornell University)|Jul 8, 2023
Machine Learning in Materials Science被引用数 5
ひとこと要約

本稿では、異なるレベルの理論を有する複数の量子力学的(QM)データセットから同時に学習できる、基礎的機械学習原子間ポテンシャル(MLIP)をメタラーニングで訓練する手法を提案する。新しい分子への迅速な適応を可能にすることで、メタラーニングは一般化性能を向上させ、誤差を低減し、ポテンシャルエネルギー面の滑らかさを向上させ、標準的な転移学習よりも薬剤様分子(例:3BPA)において優れた性能を示す。

ABSTRACT

The development of machine learning models has led to an abundance of datasets containing quantum mechanical (QM) calculations for molecular and material systems. However, traditional training methods for machine learning models are unable to leverage the plethora of data available as they require that each dataset be generated using the same QM method. Taking machine learning interatomic potentials (MLIPs) as an example, we show that meta-learning techniques, a recent advancement from the machine learning community, can be used to fit multiple levels of QM theory in the same training process. Meta-learning changes the training procedure to learn a representation that can be easily re-trained to new tasks with small amounts of data. We then demonstrate that meta-learning enables simultaneously training to multiple large organic molecule datasets. As a proof of concept, we examine the performance of a MLIP refit to a small drug-like molecule and show that pre-training potentials to multiple levels of theory with meta-learning improves performance. This difference in performance can be seen both in the reduced error and in the improved smoothness of the potential energy surface produced. We therefore show that meta-learning can utilize existing datasets with inconsistent QM levels of theory to produce models that are better at specializing to new datasets. This opens new routes for creating pre-trained, foundational models for interatomic potentials.

研究の動機と目的

  • 大規模で多様なQMデータセットを統合し、理論レベルが不一致な状態でMLIPを訓練する課題に対処すること。
  • 迅速に新しい分子系に適応可能な、転送可能で汎用的な原子間ポテンシャルを実現すること。
  • 複数のQMレベルを有するデータセットを統合する際、メタラーニングが標準的転移学習を上回ることを示すこと。
  • 既存のマルチファシリティデータセットを活用して、事前学習済みの基礎的モデルとしての原子間ポテンシャルの構築への道筋を示すこと。
  • メタラーニングを広い化学的空間にスケールさせるために、標準化されたデータフォーマットの必要性を強調すること。

提案手法

  • 異なる理論レベル(例:DFT、CCSD(T))を持つ複数のQMデータセットを同時に学習するため、メタラーニングを用いて1つのモデルを訓練する。
  • モデルがタスク間で一般化可能な初期化を学習できる二段階最適化フレームワークを採用し、少数ショットでの高速適応を可能にする。
  • QM7-x、QMugs、ANI-1x、Transition-1x、GEOMの5つの大規模有機分子データセットで事前学習を行う。
  • 異なるQMレベルのデータセット間でエネルギーを揃えるための線形スケーリングを適用し、一貫した学習信号を確保する。
  • 小さなターゲットデータセット(例:3BPA)でメタラーニング済みモデルをファインチューニングし、特化性能を評価する。
  • MLIPのベースモデルとして、メッセージパッシングニューラルネットワークアーキテクチャ(例:SchNetやPhysNet)を採用する。
Figure 1: A diverse collection of datasets, with varying levels of theory, molecule sizes, and energies, will be incorporated into a single meta-learned potential. The distributions of the number of atoms and energy of the structures contained in the datasets used for training a potential in this wo
Figure 1: A diverse collection of datasets, with varying levels of theory, molecule sizes, and energies, will be incorporated into a single meta-learned potential. The distributions of the number of atoms and energy of the structures contained in the datasets used for training a potential in this wo

実験結果

リサーチクエスチョン

  • RQ1メタラーニングは、理論レベルが不一致な複数のQMデータセットからMLIPを同時に学習可能か?
  • RQ2メタラーニングによる多様なデータセットでの事前学習は、未観測の分子系における一般化性能と精度を向上させるか?
  • RQ3ポテンシャルエネルギー面の誤差と滑らかさという観点から、メタラーニングベースの適応は標準的転移学習に比べて優れているか?
  • RQ4メタラーニングされたモデルは、最小限のファインチューニングで3BPAのような小規模で複雑な分子においてより良い性能を達成できるか?
  • RQ5QMレベルが異なるデータセットを統合する際の実用的制限は何か?また、追加のデータはいつ有益となるか?

主な発見

  • QM7-x、QMugs、ANI-1x、Transition-1x、GEOMの複数のデータセットで訓練したメタラーニングモデルは、3BPA分子に対してファインチューニングした際、標準的転移学習よりも優れた性能を示した。
  • メタラーニングモデルが生成したポテンシャルエネルギー面は、標準的転移学習よりも著しく滑らかく、ノイズが低減され物理的整合性が向上した。
  • メタラーニングモデルは3BPAテストセットでより低い平均絶対誤差(MAE)を達成し、精度と一般化性能の向上を示した。
  • 複数のデータセットでの事前学習は多様な系において性能向上をもたらしたが、ANI-1xからCCSD(T)への直接的ファインチューニングが最小の誤差を示し、特定のタスクではデータの一貫性が重要であることを示した。
  • メタラーニングにより、新たなQM計算を必要とせず、既存の多様なデータセットを効果的に活用でき、転送可能なMLIPの開発を加速できる。
  • 本研究は、メタラーニングを広い材料科学・分子科学のコミュニティにスケールさせるために、標準化されたデータフォーマットの必要性が急務であることを強調している。
Figure 2: This work uses Reptile to build a potential that incorporates information from multiple molecular datasets, calculated at different levels of theory. This meta-learned potential adapts well to new tasks, and outperforms potentials that were trained only to the data for a single task.
Figure 2: This work uses Reptile to build a potential that incorporates information from multiple molecular datasets, calculated at different levels of theory. This meta-learned potential adapts well to new tasks, and outperforms potentials that were trained only to the data for a single task.

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。