Skip to main content
QUICK REVIEW

[论文解读] Single-model uncertainty quantification in neural network potentials does not consistently outperform model ensembles

Aik Rui Tan, Shingo Urata|arXiv (Cornell University)|May 2, 2023
Machine Learning in Materials Science参考文献 73被引用 4
一句话总结

本研究评估了单模型不确定性量化(UQ)方法——均值-方差估计、深度证据回归和高斯混合模型——与神经网络集成在神经网络原子间势(NNIPs)中的表现。尽管计算成本更低,但这些单模型方法在不同数据集和指标下均未一致优于集成方法,而集成方法在主动学习循环中仍表现出更优的泛化能力和鲁棒性。

ABSTRACT

Neural networks (NNs) often assign high confidence to their predictions, even for points far out-of-distribution, making uncertainty quantification (UQ) a challenge. When they are employed to model interatomic potentials in materials systems, this problem leads to unphysical structures that disrupt simulations, or to biased statistics and dynamics that do not reflect the true physics. Differentiable UQ techniques can find new informative data and drive active learning loops for robust potentials. However, a variety of UQ techniques, including newly developed ones, exist for atomistic simulations and there are no clear guidelines for which are most effective or suitable for a given case. In this work, we examine multiple UQ schemes for improving the robustness of NN interatomic potentials (NNIPs) through active learning. In particular, we compare incumbent ensemble-based methods against strategies that use single, deterministic NNs: mean-variance estimation, deep evidential regression, and Gaussian mixture models. We explore three datasets ranging from in-domain interpolative learning to more extrapolative out-of-domain generalization challenges: rMD17, ammonia inversion, and bulk silica glass. Performance is measured across multiple metrics relating model error to uncertainty. Our experiments show that none of the methods consistently outperformed each other across the various metrics. Ensembling remained better at generalization and for NNIP robustness; MVE only proved effective for in-domain interpolation, while GMM was better out-of-domain; and evidential regression, despite its promise, was not the preferable alternative in any of the cases. More broadly, cost-effective, single deterministic models cannot yet consistently match or outperform ensembling for uncertainty quantification in NNIPs.

研究动机与目标

  • 评估单模型不确定性量化(UQ)方法是否能在训练鲁棒的神经网络原子间势(NNIPs)方面超越神经网络集成。
  • 评估多种UQ技术——均值-方差估计、深度证据回归和高斯混合模型——在不同数据区域(域内插值与域外泛化)中的有效性。
  • 确定单确定性NNs是否适合在原子模拟的主动学习中应用,其中不确定性量化对于避免非物理结构和偏差动力学至关重要。
  • 识别有前景的单模型UQ方法相较于计算成本更高的集成方法的局限性。

提出的方法

  • 本研究采用三个基准数据集:rMD17(域内插值)、氨分子翻转(中等程度外推)和体相石英玻璃(高度外推泛化)。
  • 针对每个数据集,应用多种UQ方法:均值-方差估计(MVE)、深度证据回归、高斯混合模型(GMM)和神经网络集成。
  • 通过衡量模型误差与预测不确定性之间相关性的指标来评估不确定性,涵盖多种测试集,包括对抗性采样以探测分布外性能。
  • 在主动学习循环中使用UQ得分指导数据采集,衡量各方法识别信息量丰富、高误差区域的能力。
  • GMM方法利用训练后神经网络的潜在空间表征来拟合混合模型,实现不确定性估计而无需修改网络架构。
  • 深度证据回归通过在预测输出上定义狄利克雷分布来建模不确定性,其参数通过使用合适评分规则的端到端优化进行学习。
Figure 1: Illustration of the uncertainty quantification (UQ) methods used. $x$ denotes the input to the neural networks (NNs), while $y$ is the predicted property. In the case of NNIPs, $x$ generally represents the positions and atomic numbers of the input structure, whereas $y$ is the energy and/o
Figure 1: Illustration of the uncertainty quantification (UQ) methods used. $x$ denotes the input to the neural networks (NNs), while $y$ is the predicted property. In the case of NNIPs, $x$ generally represents the positions and atomic numbers of the input structure, whereas $y$ is the energy and/o

实验结果

研究问题

  • RQ1单模型UQ方法(如均值-方差估计、深度证据回归或高斯混合模型)是否能在NNIPs的不确定性量化中始终优于神经网络集成?
  • RQ2不同UQ方法在从域内插值到高度外推泛化的不同分布偏移程度下的表现如何?
  • RQ3具有UQ能力的单确定性NNs在原子模拟的主动学习循环中,能在多大程度上实现与集成方法相当的鲁棒性?
  • RQ4深度证据回归的数值稳定性和泛化能力是否足以使其成为NNIP训练中集成方法的可行替代方案?
  • RQ5UQ方法在识别分布外构型的不确定性方面表现如何,特别是在石英玻璃等复杂体系中?

主要发现

  • 神经网络集成在所有数据集和评估指标中始终优于所有单模型UQ方法,尤其在主动学习过程中的泛化能力和鲁棒性方面表现更优。
  • 均值-方差估计(MVE)仅在域内插值任务中表现良好,无法系统识别分布外区域的高不确定性区域。
  • 高斯混合模型(GMM)在域外泛化方面表现更优,尤其在石英玻璃数据集中,但在插值区域表现较差。
  • 尽管深度证据回归在理论上具有优势,但在优化过程中表现出数值不稳定性,且在任何指标或数据集中均未超越集成方法。
  • 没有任何一种单模型UQ方法能匹配集成方法的泛化能力,表明目前计算成本更低的确定性模型尚无法在鲁棒NNIP训练中稳定替代集成方法。
  • 本研究证实,尽管计算成本较高,基于集成的UQ仍是通过对抗性采样识别分布外构型的最可靠方法。
Figure 2: (a) Hexbin plots showing (predicted) uncertainties versus squared errors of atomic forces in all molecules of all 5-fold test sets from the rMD17 data set for each considered UQ method. Note that since the uncertainties (-NLL) for the GMM method contain negative values, all uncertainties a
Figure 2: (a) Hexbin plots showing (predicted) uncertainties versus squared errors of atomic forces in all molecules of all 5-fold test sets from the rMD17 data set for each considered UQ method. Note that since the uncertainties (-NLL) for the GMM method contain negative values, all uncertainties a

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。