Skip to main content
QUICK REVIEW

[논문 리뷰] GemNet-OC: Developing Graph Neural Networks for Large and Diverse Molecular Simulation Datasets

Johannes Gasteiger, Muhammed Shuaibi|arXiv (Cornell University)|2022. 04. 06.
Machine Learning in Materials Science인용 수 54
한 줄 요약

GemNet-OC가 OC20에서 최첨단 성과를 달성하는 동시에 학습 속도는 약 10배 빠르며, 본 논문은 데이터셋 규모, 다양성, 도메인 시프트가 GNN 설계에 미치는 영향을 분석하고, 효율적 개발을 위한 대표적인 하위집합으로 OC-2M을 도입한다.

ABSTRACT

Recent years have seen the advent of molecular simulation datasets that are orders of magnitude larger and more diverse. These new datasets differ substantially in four aspects of complexity: 1. Chemical diversity (number of different elements), 2. system size (number of atoms per sample), 3. dataset size (number of data samples), and 4. domain shift (similarity of the training and test set). Despite these large differences, benchmarks on small and narrow datasets remain the predominant method of demonstrating progress in graph neural networks (GNNs) for molecular simulation, likely due to cheaper training compute requirements. This raises the question -- does GNN progress on small and narrow datasets translate to these more complex datasets? This work investigates this question by first developing the GemNet-OC model based on the large Open Catalyst 2020 (OC20) dataset. GemNet-OC outperforms the previous state-of-the-art on OC20 by 16% while reducing training time by a factor of 10. We then compare the impact of 18 model components and hyperparameter choices on performance in multiple datasets. We find that the resulting model would be drastically different depending on the dataset used for making model choices. To isolate the source of this discrepancy we study six subsets of the OC20 dataset that individually test each of the above-mentioned four dataset aspects. We find that results on the OC-2M subset correlate well with the full OC20 dataset while being substantially cheaper to train on. Our findings challenge the common practice of developing GNNs solely on small datasets, but highlight ways of achieving fast development cycles and generalizable results via moderately-sized, representative datasets such as OC-2M and efficient models such as GemNet-OC. Our code and pretrained model weights are open-sourced.

연구 동기 및 목표

  • 작은 데이터셋에서의 GNN 개선이 크고 다양한 분자 데이터로 일반화되는지 조사한다.

제안 방법

  • OC20를 기반으로 한 멀티-레벨 상호작용 계층 및 엣지/원자 임베딩을 갖춘 GemNet-OC 기준모형을 개발한다
  • 그래프 구축을 위한 고정된 거리 컷오프를 이웃의 고정된 수로 대체한다
  • 방사/각도 항의 계산 비용을 줄이기 위해 기저 함수들을 단순화하고 최적화한다
  • 원자-원자, 엣지-원자, 엣지-간 상호작용을 포함하는 계층적 상호작용 체계를 도입하고 장거리 원자 수준 경로를 도입한다
  • 상호작용 블록에서 임베딩을 출력하고 연결해 최종 에너지/힘 예측을 개선한다
  • 재현 가능한 개발을 위해 OC20 및 OC-2M에서의 오픈 소스 코드와 사전학습 가중치를 제공한다.

실험 결과

연구 질문

  • RQ1네 가지 데이터셋 속성(화학적 다양성, 시스템 크기, 데이터셋 크기, 도메인 시프트)이 GNN 설계 결정에 어떤 영향을 미치는가?
  • RQ2OC20에서 학습된 모델이 학습 비용을 줄이며 최첨단 성능을 달성할 수 있는가?
  • RQ3OC-2M이 OC20의 추세와 잘 연관되어 더 빠른 개발을 가능하게 하는 신뢰할 수 있는 하위집합인가?
  • RQ4모델 구성요소의 효과가 작은 데이터셋과 큰 데이터셋, 그리고 OC20 하위집합 간에 다르게 나타나는가?
  • RQ5대규모 분자 데이터셋의 빠르고 확장 가능한 학습을 가능하게 하는 최적의 구조 및 학습 전략은 무엇인가?

주요 결과

  • GemNet-OC는 OC20 작업에서 최첨단 결과를 달성하며 이전의 대형 모델에 비해 학습 속도가 약 10배 빠르다
  • GemNet-OC는 OC-2M 및 OC20 학습 데이터에 대해 이전 모델들보다 우수한 성능을 보이면서도 훨씬 적은 학습 데이터를 사용한다
  • 소형 데이터셋과 대형 데이터셋에서 모델 구성요소의 성능이 현저히 차이가 있다는 점이 있다
  • OC-2M은 전체 OC20 결과와 잘 상관되어 빠르고 대표적인 모델 개발을 가능하게 한다
  • 장거리 원자 수준 경로와 최적화된 기저 함수로 구성된 상호작용 계층은 대규모 시스템을 효율적으로 처리하는 데 기여한다

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.