Skip to main content
QUICK REVIEW

[Paper Review] A recipe for scalable attention-based MLIPs: unlocking long-range accuracy with all-to-all node attention

Eric Qu, Brandon M. Wood|arXiv (Cornell University)|Mar 6, 2026
Machine Learning in Materials Science0 citations
TL;DR

The paper presents a scalable attention-based framework for machine-learned interatomic potentials (MLIPs) that uses all-to-all node attention to improve long-range interaction accuracy.

ABSTRACT

Machine-learning interatomic potentials (MLIPs) have advanced rapidly, with many top models relying on strong physics-based inductive biases. However, as models scale to larger systems like biomolecules and electrolytes, they struggle to accurately capture long-range (LR) interactions, leading current approaches to rely on explicit physics-based terms or components. In this work, we propose AllScAIP, a straightforward, attention-based, and energy-conserving MLIP model that scales to O(100 million) training samples. It addresses the long-range challenge using an all-to-all node attention component that is data-driven. Extensive ablations reveal that in low-data/small-model regimes, inductive biases improve sample efficiency. However, as data and model size scale, these benefits diminish or even reverse, while all-to-all attention remains critical for capturing LR interactions. Our model achieves state-of-the-art energy/force accuracy on molecular systems, as well as a number of physics-based evaluations (OMol25), while being competitive on materials (OMat24) and catalysts (OC20). Furthermore, it enables stable, long-timescale MD simulations that accurately recover experimental observables, including density and heat of vaporization predictions.

Motivation & Objective

  • Motivate the need for scalable attention mechanisms in MLIPs to capture long-range interactions in large systems.
  • Introduce an all-to-all node attention strategy to enable comprehensive inter-node communication within MLIPs.
  • Outline a scalable recipe that balances accuracy and computational efficiency for long-range atomic interactions.

Proposed method

  • Proposes an attention-based MLIP architecture with all-to-all node attention to model interatomic interactions.
  • Incorporates mechanisms to maintain scalability while enabling broad inter-node communication.
  • Outlines core components and training considerations to achieve long-range accuracy.
  • Discusses practical considerations for implementing scalable attention in MLIPs.

Experimental results

Research questions

  • RQ1How can all-to-all node attention be leveraged to improve long-range accuracy in MLIPs?
  • RQ2What are the scalability implications of using all-to-all attention in MLIPs for large systems?
  • RQ3What design choices balance accuracy and computational efficiency in scalable attention-based MLIPs?
  • RQ4How does the proposed recipe compare to existing approaches in capturing long-range interactions?

Key findings

  • The approach aims to improve long-range interaction accuracy in MLIPs through all-to-all node attention.
  • The paper discusses scalability strategies for applying attention mechanisms to MLIPs.
  • The proposed recipe addresses practical implementation aspects for scalable attention in interatomic potentials.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.