Skip to main content
QUICK REVIEW

[论文解读] Robust data-driven approach for predicting the configurational energy of high entropy alloys

Jiaxin Zhang, Xianglin Liu|arXiv (Cornell University)|Aug 10, 2019
High Entropy Alloys Studies参考文献 51被引用 9
一句话总结

该论文提出了一种鲁棒的贝叶斯数据驱动框架,用于基于有效成对相互作用(EPI)和集成采样,预测高熵合金(HEAs)的构型能。通过结合贝叶斯正则化回归与基于贝叶斯信息准则(BIC)的特征选择,该方法即使在第一性原理数据有限的情况下,也能实现高精度、可量化不确定性的预测,能量误差低于1 meV,且在NbMoTaW、NbMoTaWV和NbMoTaWTi HEAs中均实现了稳定的参数估计。

ABSTRACT

High entropy alloys (HEAs) have been increasingly attractive as promising next-generation materials due to their various excellent properties. It's necessary to essentially characterize the degree of chemical ordering and identify order-disorder transitions through efficient simulation and modeling of thermodynamics. In this study, a robust data-driven framework based on Bayesian approaches is proposed and demonstrated on the accurate and efficient prediction of configurational energy of high entropy alloys. The proposed effective pair interaction (EPI) model with ensemble sampling is used to map the configuration and its corresponding energy. Given limited data calculated by first-principles calculations, Bayesian regularized regression not only offers an accurate and stable prediction but also effectively quantifies the uncertainties associated with EPI parameters. Compared with the arbitrary determination of model complexity, we further conduct a physical feature selection to identify the truncation of coordination shells in EPI model using Bayesian information criterion. The results achieve efficient and robust performance in predicting the configurational energy, particularly given small data. The developed methodology is applied to study a series of refractory HEAs, i.e. NbMoTaW, NbMoTaWV and NbMoTaWTi where it is demonstrated how dataset size affects the confidence we can place in statistical estimates of configurational energy when data are sparse.

研究动机与目标

  • 开发一种鲁棒的、基于数据驱动的方法,用于在第一性原理数据有限的情况下预测高熵合金(HEAs)的构型能。
  • 通过基于统计标准的物理特征选择,解决多组分系统中的模型复杂性与过拟合问题。
  • 提高用于蒙特卡洛模拟有序-无序转变的代理哈密顿量的可靠性与准确性。
  • 量化有效成对相互作用(EPI)参数的不确定性,并评估其对数据集大小的敏感性。
  • 证明集成采样在增强复杂构型空间中数据代表性方面的有效性。

提出的方法

  • 采用有效成对相互作用(EPI)模型,基于选定数个配位壳层内的原子对相互作用来表示构型能。
  • 使用贝叶斯正则化回归拟合第一性原理DFT计算的能量,以实现不确定性量化与相关性估计。
  • 应用贝叶斯信息准则(BIC)进行物理特征选择,通过识别配位壳层的最优截断来实现。
  • 实施一种集成采样策略,结合具有不同短程与长程有序性的多种超胞,以增强数据的代表性。
  • 通过在不同大小的训练数据集上进行交叉验证,评估模型的鲁棒性与泛化性能。
  • 分析EPI参数的方差-协方差矩阵,以评估不同配位壳层间参数的相互依赖性与稳定性。

实验结果

研究问题

  • RQ1如何在第一性原理数据有限的情况下,实现对高熵合金中构型能的高精度预测?
  • RQ2对于不同大小的数据集,EPI模型中应包含多少个配位壳层为最优?
  • RQ3与任意模型截断相比,基于BIC的贝叶斯特征选择在多大程度上提升了模型的鲁棒性?
  • RQ4数据集大小在多大程度上影响EPI参数估计的不确定性和稳定性?
  • RQ5集成采样能否有效表征多组分HEAs的复杂构型空间?

主要发现

  • 所提出的贝叶斯框架在所有三种HEAs(NbMoTaW、NbMoTaWV、NbMoTaWTi)中均实现了小于1 meV/原子的均方根误差(RMSE),表明预测精度极高。
  • 通过基于BIC的特征选择,发现小数据集(nt < 100)下最优配位壳层数为m=2–3,中等数据集(100 ≤ nt ≤ 400)为m=5–6,大数据集(nt > 400)为m=6–9,有效减少了过拟合与欠拟合。
  • 随着数据集大小从nt=100增加到nt=400,EPI参数的不确定性,尤其是最近邻壳层的参数,显著降低,表明参数稳定性得到提升。
  • 方差-协方差矩阵揭示了最近邻壳层内存在非平凡的相关性,尽管假设其相互独立,凸显了考虑相关性的建模的重要性。
  • 与单一切片超胞采样相比,集成采样策略使RMSE的标准差更低(ε < 1 meV),证明了其在增强数据代表性方面的优势。
  • 在NbMoTaWV体系中,BIC识别出的m=9壳层模型优于人为选择的m=7壳层模型,表明BIC引导的选择可避免复杂体系中的欠拟合。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。