Skip to main content
QUICK REVIEW

[Paper Review] A survey of unsupervised learning methods for high-dimensional uncertainty quantification in black-box-type problems

Katiana Kontolati, Dimitrios Loukrezis|arXiv (Cornell University)|Feb 9, 2022
Probabilistic and Robust Engineering Design137 references52 citations
TL;DR

This paper proposes a manifold-based polynomial chaos expansion (m-PCE) framework that uses unsupervised dimension reduction (DR) techniques to enable efficient high-dimensional uncertainty quantification (UQ) in black-box PDE models. By projecting high-dimensional stochastic inputs onto low-dimensional manifolds via 13 DR methods—linear (e.g., PCA, ICA) and nonlinear (e.g., AE, LLE)—and then constructing sparse PCE surrogates, the approach achieves comparable accuracy to expensive deep learning surrogates with orders-of-magnitude faster training, especially for problems with low intrinsic dimensionality.

ABSTRACT

Constructing surrogate models for uncertainty quantification (UQ) on complex partial differential equations (PDEs) having inherently high-dimensional $\mathcal{O}(10^{\ge 2})$ stochastic inputs (e.g., forcing terms, boundary conditions, initial conditions) poses tremendous challenges. The curse of dimensionality can be addressed with suitable unsupervised learning techniques used as a pre-processing tool to encode inputs onto lower-dimensional subspaces while retaining its structural information and meaningful properties. In this work, we review and investigate thirteen dimension reduction methods including linear and nonlinear, spectral, blind source separation, convex and non-convex methods and utilize the resulting embeddings to construct a mapping to quantities of interest via polynomial chaos expansions (PCE). We refer to the general proposed approach as manifold PCE (m-PCE), where manifold corresponds to the latent space resulting from any of the studied dimension reduction methods. To investigate the capabilities and limitations of these methods we conduct numerical tests for three physics-based systems (treated as black-boxes) having high-dimensional stochastic inputs of varying complexity modeled as both Gaussian and non-Gaussian random fields to investigate the effect of the intrinsic dimensionality of input data. We demonstrate both the advantages and limitations of the unsupervised learning methods and we conclude that a suitable m-PCE model provides a cost-effective approach compared to alternative algorithms proposed in the literature, including recently proposed expensive deep neural network-based surrogates and can be readily applied for high-dimensional UQ in stochastic PDEs.

Motivation & Objective

  • To address the curse of dimensionality in high-dimensional uncertainty quantification (UQ) for complex PDEs with stochastic inputs.
  • To evaluate the effectiveness of 13 unsupervised dimension reduction (DR) methods in preserving input structure and enabling accurate surrogate modeling.
  • To develop and validate a manifold PCE (m-PCE) framework that combines DR with polynomial chaos expansions for efficient UQ.
  • To compare the predictive accuracy and computational cost of DR methods across diverse physics-based systems with varying input complexity.
  • To guide researchers in selecting optimal DR techniques based on problem complexity and computational constraints.

Proposed method

  • Apply 13 unsupervised DR methods—including linear (PCA, ICA, k-PCA), nonlinear (LLE, autoencoders), and spectral methods—to project high-dimensional stochastic inputs into low-dimensional latent spaces.
  • Use the resulting low-dimensional embeddings as inputs to construct sparse polynomial chaos expansions (PCEs) for surrogate modeling.
  • Construct m-PCE models by combining any DR method with PCE, where 'manifold' refers to the latent space from the DR step.
  • Train PCE surrogates using compressive sensing and basis adaptation techniques to ensure sparsity and accuracy.
  • Evaluate performance using relative error in moment estimation and CPU time across three physics-based black-box models: Poisson, heat, and Brusselator equations.
  • Use both Gaussian and non-Gaussian random fields as stochastic inputs to assess robustness across input distributions.

Experimental results

Research questions

  • RQ1Which unsupervised DR methods most effectively reduce the dimensionality of high-dimensional stochastic inputs while preserving structural and statistical properties?
  • RQ2How does the predictive accuracy of m-PCE models vary across different DR methods when applied to physics-based PDEs with complex input uncertainties?
  • RQ3What is the trade-off between computational cost (CPU time) and accuracy across linear versus nonlinear DR methods in high-dimensional UQ?
  • RQ4Can simple DR methods like PCA or ICA achieve comparable accuracy to complex deep learning-based DR methods such as autoencoders in high-dimensional UQ?
  • RQ5Under what conditions does the intrinsic dimensionality of the input data determine the success of the m-PCE framework?

Key findings

  • For simple 1D problems with low-complexity stochastic fields, linear DR methods like PCA, k-PCA, and ICA achieved the best balance of accuracy and speed, with training times in the order of seconds.
  • For complex problems involving multi-scale stochastic fields and spatiotemporal responses (e.g., heat and Brusselator equations), nonlinear DR methods such as autoencoders and LLE outperformed linear methods in accuracy, albeit with 1–3 orders of magnitude higher CPU cost.
  • The optimal PCE surrogate was consistently constructed with a maximum polynomial degree of 2–4, indicating that DR reveals a smooth, low-dimensional manifold amenable to low-order polynomial approximation.
  • The m-PCE framework with standard PCA for DR and PCE for modeling achieved predictive accuracy comparable to deep neural network-based surrogates, but with training times up to 4 orders of magnitude faster.
  • When both accuracy and computational cost are considered, simple DR methods often outperform complex, overparameterized alternatives, especially in problems with low intrinsic dimensionality.
  • The study confirms that m-PCE is a cost-effective alternative to expensive deep learning surrogates, particularly when the input data possess a low intrinsic dimensionality that can be effectively captured by unsupervised DR.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.