Skip to main content
QUICK REVIEW

[论文解读] Comparison of surrogate-based uncertainty quantification methods for computationally expensive simulators

Nathan Owen, Peter Challenor|arXiv (Cornell University)|Nov 3, 2015
Probabilistic and Robust Engineering Design参考文献 50被引用 6
一句话总结

本文针对计算成本高昂的模拟器中的不确定性量化,比较了多项式混沌(PC)与高斯过程(GP)代理模型。基于两个工业级模拟器——adJULES 和 VEGACONTROL,以及多种实验设计,研究发现:在非线性场景中,尤其是设计规模较小至中等时,GP 代理模型通常优于 PC;尽管 PC 计算成本更低,在简单情形下具有竞争力,但其在非线性问题中表现较差,尤其在输出分布为双峰时表现不佳。

ABSTRACT

Polynomial chaos and Gaussian process emulation are methods for surrogate-based uncertainty quantification, and have been developed independently in their respective communities over the last 25 years. Despite tackling similar problems in the field, to our knowledge there has yet to be a critical comparison of the two approaches in the literature. We begin by providing a detailed description of polynomial chaos and Gaussian process approaches for building a surrogate model of a black-box function. The accuracy of each surrogate method is then tested and compared for two simulators used in industry: a land-surface model (adJULES) and a launch vehicle controller (VEGACONTROL). We analyse surrogates built on experimental designs of various size and type to investigate their performance in a range of modelling scenarios. Specifically, polynomial chaos and Gaussian process surrogates are built on Sobol sequence and tensor grid designs. Their accuracy is measured by their ability to estimate the mean, standard deviation, exceedance probabilities and probability density function of the simulator output, as well as a root mean square error metric, based on an independent validation design. We find that one method does not unanimously outperform the other, but advantages can be gained in some cases, such that the preferred method depends on the modelling goals of the practitioner. Our conclusions are likely to depend somewhat on the modelling choices for the surrogates as well as the design strategy. We hope that this work will spark future comparisons of the two methods in their more advanced formulations and for different sampling strategies.

研究动机与目标

  • 为基于代理模型的不确定性量化方法,提供多项式混沌与高斯过程代理模型的批判性、直接比较。
  • 评估两种方法在不同模拟器类型、设计策略及实验设计规模下的性能表现。
  • 评估每种代理模型在估计关键不确定性指标(如均值、标准差、超越概率及概率密度函数)方面的准确性。
  • 考察每种方法的实际优势,特别是计算成本与预测不确定性可用性方面的差异。
  • 根据建模目标与问题特征,为实践者提供选择最适宜代理方法的指导。

提出的方法

  • 通过非侵入式回归方法,基于模拟器运行结果估计系数,构建多项式混沌代理模型。
  • 使用平方指数和 Matérn 核函数,结合贝叶斯推断,构建高斯过程代理模型。
  • 利用 Sobol 序列和不同规模的张量网格生成的实验设计,训练两种代理模型。
  • 采用独立的 1000 点拉丁超立方设计验证代理模型性能,用于误差度量与分布估计。
  • 通过均方根误差、均值与标准差估计误差、超越概率准确性及概率密度函数估计保真度来衡量准确性。
  • 在 GP 预测中包含 95% 置信区间,以评估其不确定性量化能力,而 PC 模型则不具备此功能。

实验结果

研究问题

  • RQ1在昂贵模拟器中,多项式混沌与高斯过程代理模型在估计均值、标准差和超越概率方面的准确性如何比较?
  • RQ2在使用小至中等规模实验设计时,哪种代理方法在非线性模拟器响应中表现更优?
  • RQ3实验设计的选择(Sobol 序列与张量网格)如何影响每种代理方法的性能?
  • RQ4在哪些场景下,一种方法能明显优于另一种,特别是在不确定性量化与计算成本方面?
  • RQ5结果在多大程度上依赖于建模选择,如多项式阶数或协方差函数的选择?

主要发现

  • 在非线性二维测试函数中,高斯过程代理模型在估计均值、标准差和超越概率方面始终优于多项式混沌代理模型。
  • 对于 adJULES 和 VEGACONTROL 模拟器,GP 代理模型在大多数设计规模和类型下表现出更高的准确性,尤其在捕捉复杂非线性行为方面表现优异。
  • 多项式混沌在更简单或低维情况下表现尚可,尤其在设计规模较大时,但在处理非线性及双峰输出分布时表现不佳。
  • 在张量网格设计下,GP 与 PC 的性能差距显著缩小,PC 在某些情况下略占优势,但 GP 仍保持更高的稳定性。
  • 高斯过程代理模型通过 95% 置信区间提供内置的预测不确定性,这是多项式混沌所不具备的关键优势,后者缺乏分布不确定性估计。
  • 尽管 GP 代理模型具有诸多优势,多项式混沌在计算成本上仍更低,因此在不确定性量化非首要目标时可能更具吸引力。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。