Skip to main content
QUICK REVIEW

[Paper Review] Distillation of atomistic foundation models across architectures and chemical domains

John L. A. Gardner, Daniel F. Thomas du Toit|ArXiv.org|Jun 12, 2025
Machine Learning in Materials Science3 citations
TL;DR

The paper presents an architecture-agnostic distillation protocol that transfers knowledge from large atomistic foundation models to smaller, faster student MLIPs via synthetic data, achieving substantial speedups (over 10x to over 100x) across diverse chemical domains. The approach enables accurate, scalable MD simulations on modest hardware by distilling into multiple architectures using a small fine-tuning set.

ABSTRACT

Machine-learned interatomic potentials have transformed computational research in the physical sciences. Recent atomistic `foundation' models have changed the field yet again: trained on many different chemical elements and domains, these potentials are widely applicable, but comparably slow and resource-intensive to run. Here we show how distillation via synthetic data can be used to cheaply transfer knowledge from atomistic foundation models to a range of different architectures, unlocking much smaller, more efficient potentials. We demonstrate speed-ups of $> 10 imes$ by distilling from one graph-network architecture into another, and $> 100 imes$ by leveraging the atomic cluster expansion framework. We showcase applicability across chemical and materials domains: from liquid water to hydrogen under extreme conditions; from porous silica and a hybrid halide perovskite solar-cell material to modelling organic reactions. Our work shows how distillation can support the routine and computationally efficient use of current and future atomistic foundation models in real-world scientific research.

Motivation & Objective

  • Demonstrate a general distillation protocol to transfer knowledge from atomistic foundation models (FMs) to smaller, faster student MLIPs across chemical domains.
  • Show architecture-agnostic applicability by distilling into multiple MLIP architectures and leveraging synthetic data labeling.
  • Quantify computational efficiency and accuracy trade-offs, including memory usage and MD stability, across representative systems.
  • Validate that distilled models preserve essential physical properties through MD-based diagnostics and benchmarks.
  • Highlight practical implications for making atomistic FMs accessible with modest hardware.

Proposed method

  • Fine-tune an existing atomistic FM on a small domain-specific structure set with quantum-mechanical labels.
  • Use the fine-tuned FM to generate a large synthetic dataset by the rattle-relax-repeat augmentation without MD simulations.
  • Train small, fast student MLIP architectures on the synthetic data to approximate the FM's predictions and labels.
  • Evaluate distilled models against DFT test-set and compare structural/thermodynamic properties in MD simulations.
  • Demonstrate speed-ups and scalability across architectures (TensorNet, PaiNN, ACE) and within the ACE/EDDP families.
  • Showcase architecture-agnostic compatibility with ASE calculators and augment-atoms to enable end-to-end workflows.

Experimental results

Research questions

  • RQ1Can synthetic-data distillation transfer knowledge from a high-capacity atomistic FM to smaller, faster student models across different architectures?
  • RQ2How much speed-up and memory efficiency can be achieved while preserving accuracy relative to DFT labels?
  • RQ3Do distilled MLIPs reproduce key structural and dynamical properties in MD across diverse chemical domains?
  • RQ4What are the practical limitations and domain boundaries of distillation for reactive and high-energy configurations?
  • RQ5How do distillation outcomes vary with architecture, cutoff radii, and amount of fine-tuning data?

Key findings

  • Distillation yields >10x speed-ups when transferring from a graph-network FM to other graph-network architectures, and >100x when leveraging the ACE framework.
  • Distilled models (TensorNet, PaiNN, ACE) achieve force MAEs close to the fine-tuned FM on DFT labels, with substantial MD-speed advantages.
  • Distilled models enable stable MD on a single GPU and scale to larger system sizes beyond the FM’s memory limits.
  • Across domains (water, hydrogen, silica, MAPI, and organic reaction in solvent), distilled models reproduce key structural and dynamical features comparable to or better than the teacher in some metrics.
  • Ablation studies show synthetic-data scaling improves FM-to-DFT accuracy, and distilled models can operate with smaller cutoffs than the FM without significant loss in accuracy.
  • The approach requires modest domain data (<50 DFT-labelled structures) for fine-tuning and is fully automated with open-source tooling.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.