[Paper Review] A Foundational Potential Energy Surface Dataset for Materials
The paper introduces MatPES, an open foundational PES dataset (~434k PBE and ~388k r2SCAN structures) and shows UMLIPs trained on MatPES outperform those trained on larger, less-accurate datasets.
Accurate potential energy surface (PES) descriptions are essential for atomistic simulations of materials. Universal machine learning interatomic potentials (UMLIPs)$^{1-3}$ offer a computationally efficient alternative to density functional theory (DFT)$^4$ for PES modeling across the periodic table. However, their accuracy today is fundamentally constrained due to a reliance on DFT relaxation data.$^{5,6}$ Here, we introduce MatPES, a foundational PES dataset comprising $\sim 400,000$ structures carefully sampled from 281 million molecular dynamics snapshots that span 16 billion atomic environments. We demonstrate that UMLIPs trained on the modestly sized MatPES dataset can rival, or even outperform, prior models trained on much larger datasets across a broad range of equilibrium, near-equilibrium, and molecular dynamics property benchmarks. We also introduce the first high-fidelity PES dataset based on the revised regularized strongly constrained and appropriately normed (r$^2$SCAN) functional$^7$ with greatly improved descriptions of interatomic bonding. The open source MatPES initiative emphasizes the importance of data quality over quantity in materials science and enables broad community-driven advancements toward more reliable, generalizable, and efficient UMLIPs for large-scale materials discovery and design.
Motivation & Objective
- Address limitations of existing PES datasets (e.g., MPRelax) that bias UMLIP accuracy and transferability.
- Create a high-quality, chemically diverse PES dataset with improved DFT descriptions (PBE and r2SCAN).
- Demonstrate that UMLIPs trained on MatPES achieve comparable or superior accuracy with fewer structures.
- Provide open-source tools and benchmarks to foster community-driven development of UMLIPs for materials discovery.
Proposed method
- Generate a comprehensive configuration space via MD sampling from 281 million structures using a pre-trained M3GNet UMLIP.
- Apply an enhanced 2DIRECT sampling to select representative structures covering structural and atomic environments.
- Compute high-fidelity single-point energies, forces, and stresses with PBE and r2SCAN using VASP on 504,811 structures.
- Train UMLIPs (M3GNet, CHGNet, TensorNet) on MatPES PBE and MatPES r2SCAN and evaluate against MPRelax and OMat24 baselines.
- Benchmark across equilibrium, near-equilibrium, and MD properties using MatCalc benchmarks (fingerprint distance, formation energy, moduli, CV, MD stability, ionic conductivity).
Experimental results
Research questions
- RQ1Can a carefully sampled, mid-sized PES dataset yield UMLIPs that rival or surpass models trained on much larger, noisier datasets?
- RQ2Does incorporating a higher-fidelity DFT functional (r2SCAN) into MatPES improve PES descriptions across elements and bonding regimes?
- RQ3Do MatPES-trained UMLIPs deliver better MD stability and more reliable dynamical properties than models trained on MPRelax/OMat24 datasets?
- RQ4What is the value of data quality versus quantity for universal MLIPs in large-scale materials discovery?
Key findings
- MatPES-trained UMLIPs outperform MPRelax- and OMat24-trained counterparts on equilibrium, near-equilibrium, and MD benchmarks.
- MatPES PBE UMLIPs achieve lower test set errors and exhibit little overfitting (training/validation/test MAEs close to each other).
- The r2SCAN-based MatPES dataset provides improved bonding descriptions and comparable or better performance across properties.
- MD stability is markedly higher for MatPES UMLIPs; fewer terminations in high-temperature MD runs compared to MPRelax/OMat24 baselines.
- Equivariant TensorNet models generally show better MD stability and conductivity predictions than invariant architectures within MatPES.
- The MatCalc benchmark suite demonstrates broad, cross-property improvements, underscoring data quality over sheer dataset size.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.