[论文解读] Convergence of uncertainty estimates in Ensemble and Bayesian sparse model discovery
本文为基于自助采样法的序列阈值最小二乘法(STLS)的集成稀疏模型发现建立了理论保证,证明了虚假发现概率和真实发现概率的指数收敛性。结果表明,该方法在计算效率上高效,且不确定性量化效果与昂贵的贝叶斯MCMC方法相当,对噪声和超参数变化具有鲁棒性。
Sparse model identification enables nonlinear dynamical system discovery from data. However, the control of false discoveries for sparse model identification is challenging, especially in the low-data and high-noise limit. In this paper, we perform a theoretical study on ensemble sparse model discovery, which shows empirical success in terms of accuracy and robustness to noise. In particular, we analyse the bootstrapping-based sequential thresholding least-squares estimator. We show that this bootstrapping-based ensembling technique can perform a provably correct variable selection procedure with an exponential convergence rate of the error rate. In addition, we show that the ensemble sparse model discovery method can perform computationally efficient uncertainty estimation, compared to expensive Bayesian uncertainty quantification methods via MCMC. We demonstrate the convergence properties and connection to uncertainty quantification in various numerical studies on synthetic sparse linear regression and sparse model discovery. The experiments on sparse linear regression support that the bootstrapping-based sequential thresholding least-squares method has better performance for sparse variable selection compared to LASSO, thresholding least-squares, and bootstrapping-based LASSO. In the sparse model discovery experiment, we show that the bootstrapping-based sequential thresholding least-squares method can provide valid uncertainty quantification, converging to a delta measure centered around the true value with increased sample sizes. Finally, we highlight the improved robustness to hyperparameter selection under shifting noise and sparsity levels of the bootstrapping-based sequential thresholding least-squares method compared to other sparse regression methods.
研究动机与目标
- 为基于自助采样法的STLS的集成稀疏模型发现建立理论基础。
- 证明该方法在虚假发现率和真实发现率上实现正确变量选择,并具有指数收敛性。
- 证明集成STLS能够实现计算高效的不确定性量化,其效果可与贝叶斯MCMC方法相媲美。
- 分析该方法在不同噪声水平和稀疏性条件下的鲁棒性。
- 通过渐近等价性,将基于自助采样法的STLS与贝叶斯稀疏推断联系起来。
提出的方法
- 该方法采用基于袋装包含概率的STLS,通过自助采样估计变量选择的稳定性。
- 对每个自助采样样本应用序列阈值最小二乘法,以识别相关特征。
- 最终模型基于各候选项在自助采样重复中的包含频率进行选择。
- 理论分析依赖于二项分布试验上的集中不等式和并集界,以推导误差率。
- 方差估计通过自助采样重复的实证分布近似得到。
- 通过对比后验分布与包含概率分布的理论比较,建立了与贝叶斯spline-and-slab先验的渐近等价性。
实验结果
研究问题
- RQ1基于自助采样法的STLS是否在虚假发现概率和真实发现概率上实现指数收敛?
- RQ2集成STLS能否提供随样本量增加而收敛于真实参数分布的不确定性估计?
- RQ3与LASSO和阈值最小二乘法相比,该方法在变量选择准确性和鲁棒性方面表现如何?
- RQ4该集成方法是否在渐近意义上等价于具有spline-and-slab先验的贝叶斯稀疏推断?
- RQ5该方法在不同噪声水平和稀疏性条件下表现如何?
主要发现
- 在比先前方法更宽松的正则性条件下,基于自助采样法的STLS估计器在虚假发现概率和真实发现概率上均实现了指数收敛。
- 在合成线性回归实验中,该方法在稀疏变量选择方面优于LASSO、阈值最小二乘法以及基于自助采样法的LASSO。
- 集成方法的不确定性估计随样本量增加而收敛至以真实模型参数为中心的狄拉克测度,表明估计具有一致性。
- 在所有活跃指标上,E-SINDy与贝叶斯SINDy估计之间的Wasserstein-2距离迅速收敛至零,证实了分布等价性。
- 与其它稀疏回归技术相比,该方法在噪声和稀疏性水平变化时表现出更强的超参数选择鲁棒性。
- 理论分析证实,该集成方法通过自助采样重复提供了有效的方差估计,支持可靠的不确定性量化。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。