Skip to main content
QUICK REVIEW

[논문 리뷰] Learning Together: Towards foundational models for machine learning interatomic potentials with meta-learning

Alice E. A. Allen, Nicholas Lubbers|arXiv (Cornell University)|2023. 07. 08.
Machine Learning in Materials Science인용 수 5
한 줄 요약

이 논문은 다양한 수준의 이론을 가진 여러 양자역학적(QM) 데이터셋에서 동시에 학습할 수 있도록 메타학습을 사용하여 기초적인 기계학습 상호작용 잠재에너지함수(MLIPs)를 훈련시키는 것을 제안한다. 새로운 분자로의 빠른 적응을 가능하게 함으로써 메타학습은 일반화 능력을 향상시키고 오차를 줄이며 잠재에너지 표면의 매끄러움을 향상시켜, 표준 전이학습에 비해 3BPA와 같은 약물 유사 분자에서 뛰어난 성능을 보인다.

ABSTRACT

The development of machine learning models has led to an abundance of datasets containing quantum mechanical (QM) calculations for molecular and material systems. However, traditional training methods for machine learning models are unable to leverage the plethora of data available as they require that each dataset be generated using the same QM method. Taking machine learning interatomic potentials (MLIPs) as an example, we show that meta-learning techniques, a recent advancement from the machine learning community, can be used to fit multiple levels of QM theory in the same training process. Meta-learning changes the training procedure to learn a representation that can be easily re-trained to new tasks with small amounts of data. We then demonstrate that meta-learning enables simultaneously training to multiple large organic molecule datasets. As a proof of concept, we examine the performance of a MLIP refit to a small drug-like molecule and show that pre-training potentials to multiple levels of theory with meta-learning improves performance. This difference in performance can be seen both in the reduced error and in the improved smoothness of the potential energy surface produced. We therefore show that meta-learning can utilize existing datasets with inconsistent QM levels of theory to produce models that are better at specializing to new datasets. This opens new routes for creating pre-trained, foundational models for interatomic potentials.

연구 동기 및 목표

  • 다양한 수준의 이론을 가진 대규모이고 이질적인 QM 데이터셋을 결합하여 MLIPs를 훈련시키는 데 도전하는 데 목적을 두다.
  • 새로운 분자 시스템에 빠르게 적응할 수 있는 이식 가능하고 일반적인 상호작용 잠재에너지함수를 가능하게 하는 것.
  • 여러 개의 QM 수준을 가진 데이터셋을 통합할 때 메타학습이 표준 전이학습보다 뛰어난 성능을 보일 수 있음을 입증하는 것.
  • 기존의 다중 정밀도 데이터셋을 활용하여 기초 모델을 사전 훈련하는 길을 마련하는 것.
  • 메타학습을 보다 넓은 화학적 공간으로 확장하기 위해 표준화된 데이터 포맷의 필요성을 부각하는 것.

제안 방법

  • 다양한 수준의 이론(예: DFT, CCSD(T))을 가진 여러 QM 데이터셋을 동시에 학습할 수 있도록 메타학습을 활용한다.
  • 모델이 작업 간 일반화가 가능한 초기화를 학습할 수 있도록 이중 최적화 프레임워크를 사용한다. 이는 소수의 샘플로도 신속한 적응을 가능하게 한다.
  • 5개의 대규모 유기 분자 데이터셋(QM7-x, QMugs, ANI-1x, Transition-1x, GEOM)에서 사전 훈련을 수행한다.
  • 다른 QM 수준을 가진 데이터셋 간의 에너지를 일치시키기 위해 선형 스케일링을 적용하여 일관된 훈련 신호를 확보한다.
  • 소규모 타겟 데이터셋(예: 3BPA)에서 메타학습된 모델을 미세조정하여 특화 성능을 평가한다.
  • 메시지 전달 신경망 아키텍처(예: SchNet 또는 PhysNet)를 MLIP의 기본 모델로 사용한다.
Figure 1: A diverse collection of datasets, with varying levels of theory, molecule sizes, and energies, will be incorporated into a single meta-learned potential. The distributions of the number of atoms and energy of the structures contained in the datasets used for training a potential in this wo
Figure 1: A diverse collection of datasets, with varying levels of theory, molecule sizes, and energies, will be incorporated into a single meta-learned potential. The distributions of the number of atoms and energy of the structures contained in the datasets used for training a potential in this wo

실험 결과

연구 질문

  • RQ1메타학습을 통해 이질적인 QM 수준을 가진 여러 QM 데이터셋에서 동시에 MLIPs를 학습시킬 수 있는가?
  • RQ2메타학습을 통해 다양한 데이터셋에서 사전 훈련한 모델이 새로운 분자 시스템에서 일반화와 정확도를 향상시키는가?
  • RQ3잠재에너지 표면의 오차와 매끄러움 측면에서 메타학습 기반 적응은 표준 전이학습보다 어떻게 다를까?
  • RQ4메타학습된 모델은 최소한의 미세조정으로도 복잡한 소규모 분자인 3BPA에서 더 뛰어난 성능을 낼 수 있는가?
  • RQ5다양한 QM 수준을 가진 데이터셋을 결합할 때의 실질적 한계는 무엇이며, 추가 데이터는 언제 유익한가?

주요 결과

  • QM7-x, QMugs, ANI-1x, Transition-1x, GEOM 등 여러 데이터셋에서 훈련한 메타학습 모델은 3BPA 분자에서 표준 전이학습에 비해 향상된 성능을 보였다.
  • 메타학습 모델이 생성한 잠재에너지 표면은 표준 전이학습보다 훨씬 매끄럽게 나타나 소음이 감소하고 물리적 일관성이 향상되었다.
  • 메타학습 모델은 3BPA 테스트 세트에서 더 낮은 평균 절대 오차(MAE)를 기록하여 정확도와 일반화 능력이 향상됨을 입증했다.
  • 다양한 시스템에서의 성능 향상을 위해 여러 데이터셋에서의 사전 훈련이 유익했지만, ANI-1x에서 CCSD(T)로 직접 미세조정한 결과가 가장 낮은 오차를 기록하여 특정 작업에서는 데이터 일관성이 중요함을 시사했다.
  • 메타학습은 새로운 QM 계산을 요구하지 않고 기존의 이질적인 데이터셋을 효과적으로 활용할 수 있어, 이식 가능한 MLIPs 개발을 가속화한다.
  • 본 연구는 메타학습을 더 넓은 재료 및 분자 과학 분야로 확장하기 위해 표준화된 데이터 포맷의 긴급한 필요성을 부각시킨다.
Figure 2: This work uses Reptile to build a potential that incorporates information from multiple molecular datasets, calculated at different levels of theory. This meta-learned potential adapts well to new tasks, and outperforms potentials that were trained only to the data for a single task.
Figure 2: This work uses Reptile to build a potential that incorporates information from multiple molecular datasets, calculated at different levels of theory. This meta-learned potential adapts well to new tasks, and outperforms potentials that were trained only to the data for a single task.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.