[Paper Review] Forces are not Enough: Benchmark and Critical Evaluation for Machine Learning Force Fields with Molecular Simulations
The paper introduces a simulation-based benchmark suite for ML force fields in MD, showing force/energy accuracy alone does not guarantee realistic trajectories; stability and observables are essential, with NequIP often performing best but at higher cost.
Molecular dynamics (MD) simulation techniques are widely used for various natural science applications. Increasingly, machine learning (ML) force field (FF) models begin to replace ab-initio simulations by predicting forces directly from atomic structures. Despite significant progress in this area, such techniques are primarily benchmarked by their force/energy prediction errors, even though the practical use case would be to produce realistic MD trajectories. We aim to fill this gap by introducing a novel benchmark suite for learned MD simulation. We curate representative MD systems, including water, organic molecules, a peptide, and materials, and design evaluation metrics corresponding to the scientific objectives of respective systems. We benchmark a collection of state-of-the-art (SOTA) ML FF models and illustrate, in particular, how the commonly benchmarked force accuracy is not well aligned with relevant simulation metrics. We demonstrate when and how selected SOTA methods fail, along with offering directions for further improvement. Specifically, we identify stability as a key metric for ML models to improve. Our benchmark suite comes with a comprehensive open-source codebase for training and simulation with ML FFs to facilitate future work.
Motivation & Objective
- Motivate evaluation of ML force fields (FFs) through MD simulations, not only force/energy prediction accuracy.
- Curate diverse MD systems (water, organic molecules, peptide, materials) and define observables-based metrics.
- Assess state-of-the-art ML FF models against simulation-based objectives to identify failure modes and gaps in current approaches.
- Provide an open-source benchmark suite to standardize simulation-based evaluation of ML FFs.
Proposed method
- Define a ML FF learning setup where energies and forces are learned from atomic configurations.
- Develop a suite of physically meaningful MD observables (RDFs, h(r), diffusivity, FES) and stability criteria.
- Benchmark multiple SOTA ML FF architectures (including NequIP, GemNet, DimeNet, etc.) on four representative systems under realistic MD protocols.
- Introduce stability thresholds to detect and exclude unstable trajectory portions from observable statistics.
- Compare models on both force prediction and simulation-based metrics to reveal misalignments between force MAE and trajectory quality.
Experimental results
Research questions
- RQ1Do current SOTA ML force fields reliably simulate diverse MD systems beyond force/energy accuracy?
- RQ2What factors (e.g., stability, observables) govern the practical usability of ML FFs in MD simulations?
- RQ3Which models balance force prediction accuracy, stability, and recovery of ensemble observables across systems?
- RQ4How does training data size influence simulation-based performance across models?
Key findings
- Force accuracy alone does not align with simulation stability or recovery of ensemble statistics in MD.
- Stability is a crucial prerequisite for practical ML FF usage and can be the bottleneck even with low force error.
- Some high-force-accuracy models frequently collapse during long simulations, while others with modest force accuracy yield better MD observables.
- DeepPot-SE often offers robust simulation performance and good observable recovery with much faster runtimes at larger data budgets, while NequIP typically delivers best overall force and simulation metrics at higher cost.
- Across systems, the best force predictors can be unstable, and stable, accurate MD statistics can emerge from models with relatively modest force MAE (e.g., certain deep Pot-based or empirical approaches).
- The benchmark demonstrates that MD17, water, alanine dipeptide, and LiPS present distinct challenges, underscoring the need for diverse, simulation-based evaluation in ML FF research.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.