[论文解读] A survey of unsupervised learning methods for high-dimensional uncertainty quantification in black-box-type problems
本论文提出了一种基于流形的多项式混沌展开(m-PCE)框架,利用无监督降维(DR)技术实现对黑箱PDE模型中高维不确定性量化(UQ)的高效处理。通过将高维随机输入投影到低维流形上(采用13种DR方法——线性方法(如PCA、ICA)和非线性方法(如AE、LLE)),并构建稀疏PCE代理模型,该方法在训练速度上比昂贵的深度学习代理模型快数个数量级,同时保持相当的精度,尤其适用于内在维度较低的问题。
Constructing surrogate models for uncertainty quantification (UQ) on complex partial differential equations (PDEs) having inherently high-dimensional $\mathcal{O}(10^{\ge 2})$ stochastic inputs (e.g., forcing terms, boundary conditions, initial conditions) poses tremendous challenges. The curse of dimensionality can be addressed with suitable unsupervised learning techniques used as a pre-processing tool to encode inputs onto lower-dimensional subspaces while retaining its structural information and meaningful properties. In this work, we review and investigate thirteen dimension reduction methods including linear and nonlinear, spectral, blind source separation, convex and non-convex methods and utilize the resulting embeddings to construct a mapping to quantities of interest via polynomial chaos expansions (PCE). We refer to the general proposed approach as manifold PCE (m-PCE), where manifold corresponds to the latent space resulting from any of the studied dimension reduction methods. To investigate the capabilities and limitations of these methods we conduct numerical tests for three physics-based systems (treated as black-boxes) having high-dimensional stochastic inputs of varying complexity modeled as both Gaussian and non-Gaussian random fields to investigate the effect of the intrinsic dimensionality of input data. We demonstrate both the advantages and limitations of the unsupervised learning methods and we conclude that a suitable m-PCE model provides a cost-effective approach compared to alternative algorithms proposed in the literature, including recently proposed expensive deep neural network-based surrogates and can be readily applied for high-dimensional UQ in stochastic PDEs.
研究动机与目标
- 解决复杂PDE模型在随机输入下高维不确定性量化(UQ)中的维度灾难问题。
- 评估13种无监督降维(DR)方法在保持输入结构和实现精确代理建模方面的有效性。
- 开发并验证一种结合DR与多项式混沌展开的流形PCE(m-PCE)框架,以实现高效的UQ。
- 比较不同物理系统中DR方法在预测精度和计算成本方面的表现,这些系统具有不同的输入复杂度。
- 根据问题复杂度和计算约束,为研究人员提供选择最优DR技术的指导。
提出的方法
- 应用13种无监督DR方法——包括线性方法(PCA、ICA、k-PCA)、非线性方法(LLE、自编码器)以及谱方法——将高维随机输入投影到低维潜在空间。
- 将所得的低维嵌入作为输入,构建稀疏多项式混沌展开(PCE)代理模型。
- 通过将任意DR方法与PCE结合构建m-PCE模型,其中“流形”指DR步骤生成的潜在空间。
- 利用压缩感知和基底自适应技术训练PCE代理模型,以确保稀疏性和精度。
- 通过三个基于物理的黑箱模型(泊松方程、热传导方程和布鲁塞尔ator方程)的矩估计相对误差和CPU时间评估性能。
- 采用高斯和非高斯随机场作为随机输入,以评估在不同输入分布下的鲁棒性。
实验结果
研究问题
- RQ1哪些无监督DR方法在降低高维随机输入维度的同时,最有效地保持其结构和统计特性?
- RQ2在具有复杂输入不确定性的物理基PDE中,m-PCE模型的预测精度如何随DR方法的不同而变化?
- RQ3在高维UQ中,线性与非线性DR方法在计算成本(CPU时间)与精度之间的权衡如何?
- RQ4在高维UQ中,简单的DR方法(如PCA或ICA)是否能与基于深度学习的复杂DR方法(如自编码器)达到相当的精度?
- RQ5在何种条件下,输入数据的内在维度决定了m-PCE框架的成功?
主要发现
- 对于具有低复杂度随机场的简单1D问题,线性DR方法(如PCA、k-PCA、ICA)在精度与速度之间实现了最佳平衡,训练时间在秒级量级。
- 对于涉及多尺度随机场和时空响应的复杂问题(如热传导方程和布鲁塞尔ator方程),非线性DR方法(如自编码器和LLE)在精度上优于线性方法,尽管其CPU成本高出1至3个数量级。
- 最优PCE代理模型始终使用最高2至4次多项式阶次,表明DR揭示了一个平滑的低维流形,适合用低阶多项式近似。
- 采用标准PCA进行DR、PCE进行建模的m-PCE框架,其预测精度与基于深度神经网络的代理模型相当,但训练时间快达4个数量级。
- 在同时考虑精度和计算成本时,简单DR方法通常优于复杂且过参数化的替代方法,尤其在内在维度较低的问题中表现更优。
- 本研究证实,m-PCE是昂贵深度学习代理模型的高性价比替代方案,尤其当输入数据具有可被无监督DR有效捕捉的低内在维度时。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。