[Paper Review] Sensitivity analysis from a single input/output sample
This paper introduces a novel kernel-based estimator for Sobol’ indices using a single input/output sample, leveraging nearest-neighbor density estimation and kernel smoothing to achieve asymptotic normality and consistent variance estimation. The key contribution is a central limit theorem for the estimator under minimal moment assumptions, enabling reliable sensitivity analysis with minimal computational overhead.
The main objective of this paper is to estimate optimally Sobol' indices at any order when a unique input/output i.i.d.\ sample is available. Our approach stands on three main ingredients: semi-parametric estimation theory, high-order kernel estimation (inspired by the paper of Doksum in 1995), and mirror-type transformations as introduced in Bertin 2020 and Pujol 2022. We propose two different estimators. We prove that these estimators are asymptotically normal and efficient. Furthermore, we illustrate their numerical properties on standard examples.
Motivation & Objective
- To develop a sensitivity analysis estimator for Sobol’ indices that requires only a single input/output sample, avoiding costly repeated model evaluations.
- To establish theoretical guarantees—specifically a central limit theorem—under minimal moment conditions (e.g., $\mathbb{E}[Y^4] < \infty$), ensuring asymptotic normality of the estimator.
- To enable efficient estimation of first- and total-order Sobol’ indices without requiring nested or replicated designs, reducing computational cost.
- To extend existing nearest-neighbor and kernel-based methods to handle multivariate inputs and unknown input densities, improving robustness in high-dimensional settings.
Proposed method
- The method constructs a kernel-based estimator of $\mathbb{E}[\mathbb{E}[Y|X]^2]$ using a single sample of input-output pairs, where kernel smoothing is applied to the input variables to estimate conditional expectations.
- It employs a product kernel structure with bandwidth selection to balance bias and variance, ensuring consistency and asymptotic normality.
- The estimator uses a reweighting scheme based on nearest-neighbor density estimates to correct for sampling variability, replacing the need for separate design points or model replications.
- A central limit theorem is derived under the assumption $\mathbb{E}[Y^4] < \infty$, proving that the estimator converges in distribution to a normal law at rate $\sqrt{n}$.
- The method handles unknown input densities by estimating the density $f_X$ nonparametrically via kernel density estimation, allowing application in settings where the input distribution is not known a priori.
- Theoretical justification includes the construction of kernels of order $k$ in $d$ dimensions via tensorization of univariate orthogonal polynomial-based kernels, ensuring moment conditions are satisfied.
Experimental results
Research questions
- RQ1Can a single input/output sample be used to consistently estimate Sobol’ indices with theoretical guarantees?
- RQ2Does the proposed kernel-based estimator achieve asymptotic normality under minimal moment assumptions?
- RQ3How does the estimator perform in terms of bias and variance when the input density is unknown or estimated?
- RQ4Can the method efficiently estimate both first-order and total Sobol’ indices without requiring replicated or nested designs?
- RQ5What is the impact of bandwidth selection and kernel order on the convergence rate and finite-sample performance?
Key findings
- The proposed estimator achieves asymptotic normality under the minimal assumption $\mathbb{E}[Y^4] < \infty$, ensuring reliable inference for sensitivity indices.
- The central limit theorem is established for the estimator of $\mathbb{E}[\mathbb{E}[Y|X]^2]$, which is the core component of Sobol’ indices, enabling confidence interval construction.
- The method requires only a single sample of input-output pairs, eliminating the need for repeated model evaluations or complex design schemes like Pick-Freeze.
- The estimator remains consistent and asymptotically normal even when the input density $f_X$ is unknown and estimated nonparametrically, with convergence of the density estimator controlled via bandwidth selection.
- Theoretical analysis shows that the bias of the estimator is negligible for dimensions $d \leq 3$, and the method maintains good finite-sample performance through careful bandwidth and kernel design.
- A constructive method is provided for building multivariate kernels of arbitrary order $k$ in $d$ dimensions via tensorization of univariate orthogonal polynomial-based kernels, ensuring moment conditions are met.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.