[논문 리뷰] Open Materials 2024 (OMat24) Inorganic Materials Dataset and Models
저자들은 Open Materials 2024 (OMat24) 대규모 개방 DFT 데이터셋과 사전 학습된 EquiformerV2 모델을 공개하고, OMat24에서의 사전 학습 및 관련 데이터셋에서의 파인튜닝 후 MatBench Discovery에서 최첨단 성능을 입증한다.
The ability to discover new materials with desirable properties is critical for numerous applications from helping mitigate climate change to advances in next generation computing hardware. AI has the potential to accelerate materials discovery and design by more effectively exploring the chemical space compared to other computational methods or by trial-and-error. While substantial progress has been made on AI for materials data, benchmarks, and models, a barrier that has emerged is the lack of publicly available training data and open pre-trained models. To address this, we present a Meta FAIR release of the Open Materials 2024 (OMat24) large-scale open dataset and an accompanying set of pre-trained models. OMat24 contains over 110 million density functional theory (DFT) calculations focused on structural and compositional diversity. Our EquiformerV2 models achieve state-of-the-art performance on the Matbench Discovery leaderboard and are capable of predicting ground-state stability and formation energies to an F1 score above 0.9 and an accuracy of 20 meV/atom, respectively. We explore the impact of model size, auxiliary denoising objectives, and fine-tuning on performance across a range of datasets including OMat24, MPtraj, and Alexandria. The open release of the OMat24 dataset and models enables the research community to build upon our efforts and drive further advancements in AI-assisted materials science.
연구 동기 및 목표
- 오픈형 대규모 오픈 데이터와 모델을 통해 AI 기반 무기 재료 발견을 가속화한다.
- 다양한 비평형 구성에서 118M-구조의 공개 DFT 데이터셋을 공개적으로 접근 가능하게 제공한다.
- OMat24에서 사전 학습된 EquiformerV2 모델을 train하고 MatBench Discovery에서 평가한다.
- OMat24에서의 사전 학습 및 MPtrj와 Alexandria 서브셋에서의 파인튜닝을 통해 전달 학습을 평가한다.
- 오픈 코드, 데이터 및 체크포인트를 통해 재현성과 커뮤니티 기반 개선을 촉진한다.
제안 방법
- 무기 재료의 118백만 건 규모에 이르는 단일점 DFT, 휴식, MD 궤적을 포함하는 대규모 개방 데이터셋(OMat24)을 구성한다.
- Alexandria 이완 구조에서 시작하여 Boltzmann-rattling, AIMD, rattled relaxations의 세 가지 구조 생성 전략을 사용한다.
- 여러 모델 크기(S, M, L)로 OMat24에서 Graph Neural Network인 EquiformerV2를 사전 학습하고, 선택적으로 DeNS 잡음 제거 증강을 추가한다.
- MPtrj 및/또는 sAlexandria에서 미리 학습된 모델을 파인 튜닝하여 MatBench Discovery 메트릭스를 최적화한다.
- MatBench Discovery 벤치마크를 사용하여 바닥 상태 안정성 및 혹은 대합도(energy above hull)에 초점을 맞추고 F1, MAE 및 관련 지표를 보고한다.
- 훈련 데이터(CC 4.0), 코드 및 모델 가중치를 관대 한 라이선스 하에 공개한다.
실험 결과
연구 질문
- RQ1대규모의 다양한 공개 DFT 데이터셋(OMat24)에서의 사전 학습이 다운스트림 재료 발견 성능에 어떤 영향을 미치는가?
- RQ2EquiformerV2의 모델 크기와 잡음 제거 증강이 무기 재료의 성능에 어떤 영향을 주는가?
- RQ3OMat24와 OC20 데이터셋의 전달 학습이 MPtrj 및 Alexandria에서 파인튜닝한 후 MatBench Discovery 결과를 개선할 수 있는가?
- RQ4개발된 모델이 준수 벤치마크(MPtrj만 사용) 대비 비준수(다중 데이터셋) 벤치마크에서 어떻게 성능을 발휘하는가?
- RQ5다른 DFT 데이터셋(MP, WBM 등)과 함께 OMat24를 사용할 때의 한계와 고려사항은 무엇인가?
주요 결과
- OMat24 사전 학습은 상당한 이득을 제공하며, 비준수 모델의 경우 MatBench Discovery에서 에너지 MAE가 20 meV/atom에 도달했다.
- OMat24에서 사전 학습하고 MPtrj 및 sAlexandria에서 파인튜닝한 비준수 모델은 MatBench Discovery에서 F1 점수 0.916에 도달했다.
- MPtrj에서만 학습한 준수 모델은 DeNS를 사용하여 F1이 최대 0.823에 도달하고, 가장 작은 모델도 매우 효과적일 수 있다(F1 0.823).
- EquiformerV2 모델은 OMat24에서만 학습된 경우 검증/테스트 분할에서 에너지 MAE가 약 9–11 meV/atom 수준으로 나타나며, 다양성으로 인해 전체 WBM-테스트 결과는 일반적으로 더 좋지 않다.
- 잡음 제거(DeNS)는 더 작고 MPtrj-전용 데이터셋에서 성능을 향상시키지만, 크고 다양한 OMat24 데이터셋에서의 학습 시에는 영향이 덜하며 OC20에서의 전달도 파인튜닝 후 강력한 결과를 얻는다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.