[Paper Review] Orb-v3: atomistic simulation at scale
Orb-v3 introduces a universal, scalable set of interatomic potentials that greatly improves speed and memory while maintaining high accuracy; non-conservative direct models can match or exceed state-of-the-art performance on phonon and property benchmarks.
We introduce Orb-v3, the next generation of the Orb family of universal interatomic potentials. Models in this family expand the performance-speed-memory Pareto frontier, offering near SoTA performance across a range of evaluations with a >10x reduction in latency and > 8x reduction in memory. Our experiments systematically traverse this frontier, charting the trade-off induced by roto-equivariance, conservatism and graph sparsity. Contrary to recent literature, we find that non-equivariant, non-conservative architectures can accurately model physical properties, including those which require higher-order derivatives of the potential energy surface. This model release is guided by the principle that the most valuable foundation models for atomic simulation will excel on all fronts: accuracy, latency and system size scalability. The reward for doing so is a new era of computational chemistry driven by high-throughput and mesoscale all-atom simulations.
Motivation & Objective
- Motivate the need for universal MLIPs that combine accuracy with large-scale, high-throughput simulations.
- Present the Orb-v3 family as a scalable, near-SOTA solution along the performance-speed-memory Pareto frontier.
- Investigate how design choices (conservatism, neighbor limits, datasets) affect accuracy and scalability.
- Showcase techniques to improve rotational invariance and provide uncertainty estimates for practical MD workflows.
Proposed method
- Propose Orb-v3 as a family of models with the same base architecture as Orb-v2 but compiled and scaled for speed and memory efficiency.
- Use both direct (non-conservative) and conservative force/potential formulations to explore the Pareto frontier.
- Incorporate equigrad, a gradient-based regularization to induce roto-equivariance in conservative models.
- Train on large datasets (OMat24 AIMD subset; mpa for compatibility) and apply distillation from conservative to direct models to improve higher-order derivative accuracy.
- Implement efficient graph construction, limited neighbor counts, and GPU-accelerated nearest-neighbor routines to achieve high throughput.
- Provide an intrinsic per-atom confidence head to predict force errors for active learning and filtering.
![Figure 1 : The Pareto frontier for a range of universal Machine Learning Interatomic Potentials. The $K_{SRME}$ metric assesses a model’s ability to predict thermal conductivity via the Wigner formulation of heat transport [ 31 ] and requires accurate geometry optimizations as well as second and thi](https://ar5iv.labs.arxiv.org/html/2504.06231/assets/x1.png)
Experimental results
Research questions
- RQ1Can non-equivariant, non-conservative architectures achieve competitive thermodynamic and phonon-related predictions while delivering superior speed and memory efficiency?
- RQ2How do conservatism, neighbor limits, and training datasets influence the accuracy and scalability of universal MLIPs across diverse properties?
- RQ3What are the trade-offs between geometry optimization, higher-order derivative accuracy, and MD stability in Orb-v3 models?
- RQ4Can equigrad regularization improve rotational invariance and downstream symmetry-based workflows without sacrificing performance?
Key findings
- Orb-v3-direct-20-omat and Orb-v3-direct-inf-omat achieve strong speed/memory gains, enabling hundreds of forward passes per second and sub-second per large system scales.
- Non-conservative, direct models can match or exceed state-of-the-art performance on several benchmarks, including phonon and thermal conductivity tasks.
- Equigrad regularization significantly improves rotational invariance and robustness of symmetry-guided workflows.
- Conservative-inf-omat models attain the highest accuracy across multiple physical-property benchmarks while maintaining faster runtimes than many competitors.
- Orb-v3 models show excellent scalability, with direct-20-omat needing 32.8 GB GPU memory for 100k-atom tests and completing under 0.5 s, highlighting a step-change in throughput.
- Intrinsic per-atom confidence estimates correlate with force errors and can aid active learning and data curation.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.