Skip to main content
QUICK REVIEW

[Paper Review] Accuracy and Efficiency Benchmarks of Pretrained Machine Learning Potentials for Molecular Simulations

Peter Eastman, Evan Pretti|arXiv (Cornell University)|Jan 22, 2026
Machine Learning in Materials Science0 citations
TL;DR

This paper benchmarks 15 pretrained MLIPs for molecular simulations, evaluating accuracy, speed, memory, and stability to guide model selection.

ABSTRACT

The rapid development of pretrained Machine Learning Interatomic Potentials (MLIPs) that cover a wide range of molecular species has made it challenging to select the best model for a given application. We benchmark 15 pretrained MLIPs, evaluating each one on accuracy, speed, memory use, and ability to produce stable simulations. This provides an objective basis for practitioners to select the most appropriate MLIP for their own simulations, and offers insight into which factors most strongly influence model accuracy. We find that the number of model parameters and the size of the training set are both strongly correlated with accuracy, while training on charged molecules and including explicit Coulomb energy terms are less essential than one might expect. Speed and memory use are determined as much by the model architecture as by the size of the model.

Motivation & Objective

  • Motivate the need to select appropriate pretrained MLIPs for diverse molecular simulation tasks.
  • Provide an objective, standardized benchmark across multiple models.
  • Identify factors that most influence MLIP accuracy and efficiency.
  • Offer practical guidance on model selection balancing accuracy and resource use.

Proposed method

  • Benchmark 15 pretrained MLIPs across common molecular simulation tasks.
  • Evaluate accuracy, speed, memory usage, and simulation stability for each model.
  • Analyze correlations between model size, training set size, and accuracy.
  • Assess the impact of explicit Coulomb energy terms on performance.
  • Characterize how architecture choices influence speed and memory independent of parameter count.

Experimental results

Research questions

  • RQ1Which pretrained MLIPs deliver the best accuracy relative to their speed and memory costs?
  • RQ2How do model size and training set size correlate with accuracy across models?
  • RQ3Do explicit Coulomb energy terms confer any observable benefit for these MLIPs?
  • RQ4To what extent do model architecture choices determine speed and memory usage?

Key findings

  • Model accuracy correlates strongly with both the number of parameters and the training set size.
  • Explicit Coulomb energy terms do not provide a measurable benefit across the evaluated models.
  • Speed and memory usage are influenced more by model architecture than by just model size.
  • The study provides an objective basis for selecting MLIPs based on application-specific accuracy and efficiency needs.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.