Skip to main content
QUICK REVIEW

[论文解读] Reproducibility in high-throughput density functional theory: a comparison of AFLOW, Materials Project, and OQMD

Vinay I. Hegde, Christopher K. H. Borg|arXiv (Cornell University)|Jul 4, 2020
Machine Learning in Materials Science被引用 5
一句话总结

本研究通过使用相同的晶体结构,对比了 AFLOW、Materials Project 和 OQMD 三个高通量 DFT 数据库在生成能、带隙、体积和磁化率方面的表现。研究揭示了显著的可重复性挑战,尤其在带隙和磁化率方面(磁性差异最高达 15%),其根源在于赝势、DFT+U 和参考态的差异,凸显了在 HT-DFT 工作流程中实施标准化的必要性。

ABSTRACT

A central challenge in high throughput density functional theory (HT-DFT) calculations is selecting a combination of input parameters and post-processing techniques that can be used across all materials classes, while also managing accuracy-cost tradeoffs. To investigate the effects of these parameter choices, we consolidate three large HT-DFT databases: Automatic-FLOW (AFLOW), the Materials Project (MP), and the Open Quantum Materials Database (OQMD), and compare reported properties across each pair of databases for materials calculated using the same initial crystal structure. We find that HT-DFT formation energies and volumes are generally more reproducible than band gaps and total magnetizations; for instance, a notable fraction of records disagree on whether a material is metallic (up to 7%) or magnetic (up to 15%). The variance between calculated properties is as high as 0.105 eV/atom (median relative absolute difference, or MRAD, of 6%) for formation energy, 0.65 A$^3$/atom (MRAD of 4%) for volume, 0.21 eV (MRAD of 9%) for band gap, and 0.15 $\mu_{ m B}$/formula unit (MRAD of 8%) for total magnetization, comparable to the differences between DFT and experiment. We trace some of the larger discrepancies to choices involving pseudopotentials, the DFT+U formalism, and elemental reference states, and argue that further standardization of HT-DFT would be beneficial to reproducibility.

研究动机与目标

  • 评估尽管输入结构完全相同,主要数据库之间高通量 DFT (HT-DFT) 计算的可重复性。
  • 识别生成能、体积、带隙和总磁化率等计算性质变异的主要来源。
  • 评估输入参数选择(尤其是赝势、DFT+U 和元素参考态)对性质可重复性的影响。
  • 量化不同数据库之间差异相对于实验误差和 DFT-实验差异的幅度。
  • 倡导在 HT-DFT 中采用标准化协议,以提高材料数据库之间结果的可靠性与可比性。

提出的方法

  • 整合了三个大型 HT-DFT 数据库:AFLOW、Materials Project (MP) 和 Open Quantum Materials Database (OQMD),重点关注初始晶体结构相同的材料。
  • 对三个数据库中计算的性质(生成能、体积、带隙、总磁化率)进行了两两比较。
  • 通过计算中位数相对绝对差值(MRAD)来量化差异,其中 MRAD 定义为所有材料对中 |x_i - y_i| / |x_i| 的中位数。
  • 追溯差异至具体的方法选择,包括赝势类型、DFT+U 校正的使用以及元素参考态的定义。
  • 将差异的大小与实验值及已知的 DFT-实验偏差进行比较,以明确可重复性的限制范围。
  • 使用统计分析评估不同数据库之间矛盾预测(例如金属性与绝缘性、磁性与非磁性)的频率。

实验结果

研究问题

  • RQ1当从相同晶体结构出发时,AFLOW、MP 和 OQMD 在高通量 DFT 计算的生成能、体积、带隙和总磁化率方面具有多高的可重复性?
  • RQ2HT-DFT 结果中变异的主要来源是什么,特别是关于赝势、DFT+U 和元素参考态?
  • RQ3预测的电子和磁性性质(如金属性、磁化率)之间的差异在多大程度上影响了 HT-DFT 数据库的可靠性?
  • RQ4计算性质的观测差异与 DFT 和实验之间典型偏差相比如何?
  • RQ5为实现在 HT-DFT 平台之间获得可靠且一致的结果,需要在输入参数上达到何种程度的标准化?

主要发现

  • 生成能表现出中等可重复性,中位数相对绝对差值(MRAD)为 6%,对应最大方差达 0.105 eV/atom。
  • 体积的可重复性优于带隙,MRAD 为 4%,最大方差达 0.65 ų/atom。
  • 带隙表现出高度变异性,MRAD 为 9%,最大差异达 0.21 eV,表明数据库之间一致性差。
  • 总磁化率存在显著分歧,MRAD 为 8%,高达 15% 的材料被分类为磁性或非磁性不一致。
  • 金属性预测在高达 7% 的材料中存在不一致,凸显电子结构预测的不稳定性。
  • 差异主要归因于赝势选择、DFT+U 实现方式以及元素参考态定义的差异。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。