[论文解读] Nonparametric adaptive estimation of order 1 Sobol indices in stochastic models, with an application to Epidemiology
该论文提出了一种基于变形小波的非参数自适应估计器,用于随机模型中的一阶Sobol指标,无需依赖代理模型。该方法的收敛速率取决于模型的正则性,随着样本量增加,均方误差表现出明显的拐点效应,并在流行病学中对丙型肝炎病毒(HCV)传播建模进行了应用。
Global sensitivity analysis is a set of methods aiming at quantifying the contribution of an uncertain input parameter of the model (or combination of parameters) on the variability of the response. We consider here the estimation of the Sobol indices of order 1 which are commonly-used indicators based on a decomposition of the output's variance. In a deterministic framework, when the same inputs always give the same outputs, these indices are usually estimated by replicated simulations of the model. In a stochastic framework, when the response given a set of input parameters is not unique due to randomness in the model, metamodels are often used to approximate the mean and dispersion of the response by deterministic functions. We propose a new non-parametric estimator without the need of defining a metamodel to estimate the Sobol indices of order 1. The estimator is based on warped wavelets and is adaptive in the regularity of the model. The convergence of the mean square error to zero, when the number of simulations of the model tend to infinity, is computed and an elbow effect is shown, depending on the regularity of the model. Applications in Epidemiology are carried to illustrate the use of non-parametric estimators.
研究动机与目标
- 开发一种适用于输出变异性源于内部随机性的随机模型中的一阶Sobol指标的非参数自适应估计器。
- 消除对代理模型在估计随机响应的均值和方差时的依赖。
- 建立适应未知条件期望函数正则性的估计器收敛速率。
- 通过在注射吸毒者(PWID)中丙型肝炎病毒(HCV)传播的SIR模型上的应用,展示该方法的高效性与鲁棒性。
提出的方法
- 采用变形小波展开,非参数地估计条件期望 E[Y|Xℓ],这是计算一阶Sobol指标的核心。
- 基于惩罚经验风险最小化的方法选择程序,以适应函数 hℓ(xℓ) = E[Y|Xℓ] 的未知光滑度。
- 应用数据驱动的惩罚项 pen(J),以平衡小波系数估计中的偏差与方差。
- 通过将误差分解为逼近误差、估计误差和模型选择误差三部分,推导估计器的均方误差收敛速率。
- 利用Bernstein型不等式控制经验小波系数与其真实值之间的偏差。
- 证明估计器能够适应函数 hℓ 的正则性,实现最优收敛速率(对数因子范围内)。
实验结果
研究问题
- RQ1能否在不依赖代理模型的前提下,为随机模型中的一阶Sobol指标开发一种非参数自适应估计器?
- RQ2所提估计器的收敛速率如何依赖于底层函数 hℓ(xℓ) = E[Y|Xℓ] 的正则性?
- RQ3该估计器是否能在样本量 n 的意义上实现最优收敛速率,并自适应未知光滑度?
- RQ4模型选择对估计器性能有何影响?惩罚项如何校准?
- RQ5在计算成本和随机环境下的准确性方面,该方法与现有方法相比表现如何?
主要发现
- 在正则性条件下,所提估计器的均方误差为 O(log²(n)/n³/²) 量级,其收敛速率可自适应未知的条件期望函数光滑度。
- 收敛速率中观察到拐点效应:当函数 hℓ 属于Hölder球 B(α, 2, ∞) 时,最优速率为 n⁻⁸ᵃ⁄⁽⁴ᵃ⁺¹⁾,随着 α 增大,速率从 n⁻¹ 加快至更优速率。
- 估计器性能可自适应模型的正则性,实现最优收敛速率(对数因子范围内),且无需事先知晓光滑度信息。
- 与经典蒙特卡罗估计器(需 n(p+1) 次模型调用)相比,该方法显著减少了模型模拟次数。
- 理论边界表明,模型选择惩罚项 pen(J) 足够控制过拟合并确保一致性。
- 该方法在注射吸毒者(PWID)中丙型肝炎病毒(HCV)传播的SIR模型上得到验证,展示了其在真实流行病学建模中的实际适用性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。