Skip to main content
QUICK REVIEW

[论文解读] Open Materials 2024 (OMat24) Inorganic Materials Dataset and Models

Luis Barroso-Luque, Muhammed Shuaibi|arXiv (Cornell University)|Oct 16, 2024
Machine Learning in Materials Science被引用 79
一句话总结

作者发布 Open Materials 2024 (OMat24) 大规模开放 DFT 数据集以及预训练的 EquiformerV2 模型,在对 OMat24 进行预训练并在相关数据集上进行微调后,在 MatBench Discovery 上展示了最先进的性能。

ABSTRACT

The ability to discover new materials with desirable properties is critical for numerous applications from helping mitigate climate change to advances in next generation computing hardware. AI has the potential to accelerate materials discovery and design by more effectively exploring the chemical space compared to other computational methods or by trial-and-error. While substantial progress has been made on AI for materials data, benchmarks, and models, a barrier that has emerged is the lack of publicly available training data and open pre-trained models. To address this, we present a Meta FAIR release of the Open Materials 2024 (OMat24) large-scale open dataset and an accompanying set of pre-trained models. OMat24 contains over 110 million density functional theory (DFT) calculations focused on structural and compositional diversity. Our EquiformerV2 models achieve state-of-the-art performance on the Matbench Discovery leaderboard and are capable of predicting ground-state stability and formation energies to an F1 score above 0.9 and an accuracy of 20 meV/atom, respectively. We explore the impact of model size, auxiliary denoising objectives, and fine-tuning on performance across a range of datasets including OMat24, MPtraj, and Alexandria. The open release of the OMat24 dataset and models enables the research community to build upon our efforts and drive further advancements in AI-assisted materials science.

研究动机与目标

  • 推动开放、 大规模的开放数据与模型,以加速基于 AI 的无机材料发现。
  • 提供一个公开可获取的 118M 结构的 DFT 数据集,涵盖多样化的非平衡构型。
  • 训练并发布在 OMat24 上进行预训练的 EquiformerV2 模型,并在 MatBench Discovery 上评估。
  • 评估迁移学习:在 OMat24 上进行预训练并在 MPtrj 与 Alexandria 子集上微调。
  • 通过开放的代码、数据和检查点,促进可重复性和社区驱动的改进。

提出的方法

  • 构建一个大规模开放数据集(OMat24),包含 ~118 million 单点 DFT、松弛和 MD 轨迹,用于无机体块体材料。
  • 从 Alexandria 松弛结构出发,使用三种结构生成策略(Boltzmann-rattling、AIMD、 rattled relaxations)。
  • 在 OMat24 上对 EquiformerV2 图神经网络进行预训练,模型规模有 S、M、L,并可选添加 DeNS 去噪增强。
  • 在 MPtrj 和/或 sAlexandria 上对预训练模型进行微调,以优化 MatBench Discovery 指标。
  • 使用 MatBench Discovery 基准评估,重点关注基态稳定性和能量高于 hull,报告 F1、MAE 等相关指标。
  • 以开放许可证发布训练数据(CC 4.0)、代码和模型权重。

实验结果

研究问题

  • RQ1在大规模、开放且多样化的 DFT 数据集(OMat24)进行预训练,对下游材料发现性能有何影响?
  • RQ2模型规模和去噪增强对无机材料的 EquiformerV2 性能有何影响?
  • RQ3在 MPtrj 和 Alexandria 上微调后,使用 OMat24 与 OC20 数据集进行迁移学习能否改善 MatBench Discovery 的结果?
  • RQ4对比合规(仅 MPtrj)与非合规(多数据集)基准时,开发的模型表现如何?
  • RQ5在与其他 DFT 数据集(如 MP、WBM)并用训练时,使用 OMat24 的局限性与注意事项有哪些?

主要发现

  • OMat24 预训练带来显著收益,在非合规模型上在 MatBench Discovery 的能量 MAE 达到 20 meV/atom。
  • 在 OMat24 上预训练并在 MPtrj 和 sAlexandria 上微调的非合规模型,在 MatBench Discovery 上的 F1 得分为 0.916。
  • 仅在 MPtrj 上训练的合规模型,使用 DeNS 可达到 F1 高达 0.823,且最小模型也非常有效(F1 0.823)。
  • 仅在 OMat24 上训练的 EquiformerV2 模型,在验证/测试集上的能量 MAE 约 9–11 meV/atom,且由于多样性,WBM 测试结果总体较差。
  • 去噪(DeNS)对较小的、仅 MPtrj 数据集的模型有提升,但在大型、多样的 OMat24 数据集上训练时效果较小;从 OC20 的迁移在微调后也能获得强结果。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。