Skip to main content
QUICK REVIEW

[论文解读] DPA-2: a large atomic model as a multi-task learner

Duo Zhang, Xinzijian Liu|arXiv (Cornell University)|Dec 24, 2023
Machine Learning in Materials Science被引用 8
一句话总结

tldr: DPA-2 introduces a universal large atomic model (LAM) trained multi-task across diverse DFT-labeled datasets, enabling fine-tuning and distillation for efficient downstream molecular and material simulations.

ABSTRACT

The rapid advancements in artificial intelligence (AI) are catalyzing transformative changes in atomic modeling, simulation, and design. AI-driven potential energy models have demonstrated the capability to conduct large-scale, long-duration simulations with the accuracy of ab initio electronic structure methods. However, the model generation process remains a bottleneck for large-scale applications. We propose a shift towards a model-centric ecosystem, wherein a large atomic model (LAM), pre-trained across multiple disciplines, can be efficiently fine-tuned and distilled for various downstream tasks, thereby establishing a new framework for molecular modeling. In this study, we introduce the DPA-2 architecture as a prototype for LAMs. Pre-trained on a diverse array of chemical and materials systems using a multi-task approach, DPA-2 demonstrates superior generalization capabilities across multiple downstream tasks compared to the traditional single-task pre-training and fine-tuning methodologies. Our approach sets the stage for the development and broad application of LAMs in molecular and materials simulation research.

研究动机与目标

  • Motivate the need for a universal large atomic model (LAM) that generalizes across chemical and configurational spaces.
  • Propose the DPA-2 architecture and a multi-task pre-training pipeline to learn a unified chemical/configurational descriptor.
  • Develop a fine-tuning and distillation workflow to adapt the pre-trained model to downstream PES tasks with data efficiency.
  • Demonstrate improved zero-shot generalization and downstream sample efficiency compared to single-task and existing models.

提出的方法

  • Introduce a unified DPA-2 descriptor comprising repinit and repformer to create a symmetry-respecting representation.
  • Train the descriptor with a multi-task scheme over heterogeneous DFT-labeled datasets (different functionals, bases, etc.).
  • Use a set of energy/force heads per pre-training task connected to the shared descriptor, enabling task-specific fitting nets.
  • Fine-tune the pre-trained descriptor with downstream datasets, optionally reinitializing or reusing fitting nets.
  • Apply model distillation to create a faster student (e.g., DPA-1) guided by a teacher MD-guided labeling loop, iterating until target accuracy is reached.

实验结果

研究问题

  • RQ1Can a multi-task pre-trained large atomic model generalize to unseen downstream tasks with zero-shot performance close to task-specific models?
  • RQ2Does pre-training on heterogeneous DFT-labeled datasets improve robustness and generalization across alloys, compounds, and molecular systems compared to single-task pre-training?
  • RQ3How does fine-tuning data efficiency compare between pre-trained (multi-task) and scratch-trained models for downstream PES tasks?
  • RQ4What is the impact of distillation on achieving MD-ready speed while preserving accuracy for downstream simulations?

主要发现

  • Multi-task pre-training substantially enhances zero-shot generalization on downstream tasks (e.g., SemiCond-D shows notable RMSE improvements over single-task).
  • DPA-2 achieves competitive or superior accuracy in single-task benchmarks versus state-of-the-art models, and MT pre-training improves generalization across diverse datasets.
  • Fine-tuning with pre-trained descriptors reduces downstream data requirements and accelerates convergence compared to training from scratch.
  • A distillation loop yields a faster student model that retains the teacher’s accuracy, enabling efficient MD-scale simulations.
  • The framework supports a broad set of pre-training datasets ( alloys, cathodes, clusters, drugs, etc.) and downstream tasks, demonstrating generalizability across chemical spaces.

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。