Skip to main content
QUICK REVIEW

[Paper Review] Open Materials 2024 (OMat24) Inorganic Materials Dataset and Models

Luis Barroso-Luque, Muhammed Shuaibi|arXiv (Cornell University)|Oct 16, 2024
Machine Learning in Materials Science79 citations
TL;DR

The authors release the Open Materials 2024 (OMat24) large-scale open DFT dataset and pre-trained EquiformerV2 models, demonstrating state-of-the-art performance on MatBench Discovery after pre-training on OMat24 and fine-tuning on related datasets.

ABSTRACT

The ability to discover new materials with desirable properties is critical for numerous applications from helping mitigate climate change to advances in next generation computing hardware. AI has the potential to accelerate materials discovery and design by more effectively exploring the chemical space compared to other computational methods or by trial-and-error. While substantial progress has been made on AI for materials data, benchmarks, and models, a barrier that has emerged is the lack of publicly available training data and open pre-trained models. To address this, we present a Meta FAIR release of the Open Materials 2024 (OMat24) large-scale open dataset and an accompanying set of pre-trained models. OMat24 contains over 110 million density functional theory (DFT) calculations focused on structural and compositional diversity. Our EquiformerV2 models achieve state-of-the-art performance on the Matbench Discovery leaderboard and are capable of predicting ground-state stability and formation energies to an F1 score above 0.9 and an accuracy of 20 meV/atom, respectively. We explore the impact of model size, auxiliary denoising objectives, and fine-tuning on performance across a range of datasets including OMat24, MPtraj, and Alexandria. The open release of the OMat24 dataset and models enables the research community to build upon our efforts and drive further advancements in AI-assisted materials science.

Motivation & Objective

  • Motivate open, large-scale open data and models to accelerate AI-driven inorganic materials discovery.
  • Provide a Publicly accessible 118M-structure DFT dataset with diverse non-equilibrium configurations.
  • Train and release EquiformerV2 models pre-trained on OMat24 and evaluate on MatBench Discovery.
  • Assess transfer learning: pre-training on OMat24 and fine-tuning on MPtrj and Alexandria subsets.
  • Promote reproducibility and community-driven improvement through open code, data, and checkpoints.

Proposed method

  • Construct a large-scale open dataset (OMat24) with ~118 million single-point DFT, relaxations, and MD trajectories for inorganic bulk materials.
  • Use three structure-generation strategies (Boltzmann-rattling, AIMD, rattled relaxations) starting from Alexandria relaxed structures.
  • Pre-train EquiformerV2 graph neural networks on OMat24 with multiple model sizes (S, M, L) and optionally add DeNS denoising augmentation.
  • Fine-tune pre-trained models on MPtrj and/or sAlexandria to optimize MatBench Discovery metrics.
  • Evaluate using MatBench Discovery benchmarks focusing on ground-state stability and energy above hull, reporting F1, MAE, and related metrics.
  • Release training data (CC 4.0), code, and model weights under permissive licenses.

Experimental results

Research questions

  • RQ1How does pre-training on a large, diverse open DFT dataset (OMat24) impact downstream materials discovery performance?
  • RQ2What is the effect of model size and denoising augmentation on EquiformerV2 performance for inorganic materials?
  • RQ3Can transfer learning from OMat24 and OC20 datasets improve MatBench Discovery results after fine-tuning on MPtrj and Alexandria?
  • RQ4How do the developed models perform on compliant (MPtrj-only) versus non-compliant (multi-dataset) benchmarks?
  • RQ5What are the limitations and considerations when using OMat24 alongside other DFT datasets (e.g., MP, WBM) for training?

Key findings

  • OMat24 pre-training yields substantial gains, achieving an energy MAE of 20 meV/atom on MatBench Discovery for non-compliant models.
  • Non-compliant models pre-trained on OMat24 and fine-tuned on MPtrj and sAlexandria reach an F1 score of 0.916 on MatBench Discovery.
  • Compliant models trained only on MPtrj achieve F1 up to 0.823 with DeNS, and the smallest model can be highly effective (F1 0.823).
  • EquiformerV2 models trained solely on OMat24 show energy MAE around 9–11 meV/atom on validation/test splits, with overall WBM-test results generally worse due to diversity.
  • Denoising (DeNS) improves performance for smaller, MPtrj-only datasets but is less impactful when training on the large, diverse OMat24 dataset; transfer from OC20 also yields strong results after fine-tuning.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.