[论文解读] Regularization and Computation with high-dimensional spike-and-slab posterior distributions
本文提出了一种正则化的伪后验分布,记为 $̄Pi_{γ}$,其源自高维线性模型中的稀疏-密集先验。文章建立了当 $γ \downarrow 0$ 且 $p \to \infty$ 时,后验分布收缩至真实参数的充分条件,其收缩速率与真实后验一致,并提出了一种MCMC算法,其混合时间依赖于设计矩阵的相干性和初始化方式,其计算复杂度在有利条件下为 $O(pe^{s_\star^2})$。
We consider the Bayesian analysis of a high-dimensional statistical model with a spike-and-slab prior, and we study the forward-backward envelop of the posterior distribution -- denoted $\check\Pi_{\gamma}$ for some regularization parameter $\gamma>0$. Viewing $\check\Pi_\gamma$ as a pseudo-posterior distribution, we work out a set of sufficient conditions under which it contracts towards the true value of the parameter as $\gamma\downarrow 0$, and $p$ (the dimension of the parameter space) diverges to $\infty$. In linear regression models the contraction rate matches the contraction rate of the true posterior distribution. We also study a practical Markov Chain Monte Carlo (MCMC) algorithm to sample from $\check\Pi_{\gamma}$. In the particular case of the linear regression model, and focusing on models with high signal-to-noise ratios, we show that the mixing time of the MCMC algorithm depends crucially on the coherence of the design matrix, and on the initialization of the Markov chain. In the most favorable cases, we show that the computational complexity of the algorithm scales with the dimension $p$ as $O(pe^{s_\star^2})$, where $s_\star$ is the number of non-zeros components of the true parameter. We provide some simulation results to illustrate the theory. Our simulation results also suggest that the proposed algorithm (as well as a version of the Gibbs sampler of Narisetti and He (2014)) mix poorly when poorly initialized, or if the design matrix has high coherence.
研究动机与目标
- 研究从高维模型中稀疏-密集先验导出的正则化伪后验 $̄Pi_{γ}$ 的频率学性质。
- 建立 $̄Pi_{γ}$ 在维度 $p \to \infty$ 且 $γ \downarrow 0$ 时收缩至真实参数的充分条件。
- 分析针对线性回归模型中 $̄Pi_{γ}$ 的MCMC算法的混合时间与计算复杂度。
- 研究设计矩阵的相干性与初始化对高信噪比环境下MCMC收敛性的影响。
提出的方法
- 本文将 $̄Pi_{γ}$ 定义为真实稀疏-密集后验的前向-后向包络,作为正则化的伪后验。
- 推导出 $̄Pi_{γ}$ 在高维线性模型中以与真实后验相同速率收缩至真实参数的充分条件。
- 提出一种Metropolis-Hastings MCMC算法以从 $̄Pi_{γ}$ 中抽样,并对其混合时间进行了详细分析。
- 分析表明,混合时间与设计矩阵的相干性以及马氏链的初始化密切相关,尤其是在高信噪比设置下。
- 证明了MCMC算法的计算复杂度在有利条件下与 $O(pe^{s_\star^2})$ 成正比,其中 $s_\star$ 为真实参数向量中非零分量的数量。
- 理论结果通过模拟研究得到支持,评估了在不同设计相干性和初始化策略下MCMC的混合性能。
实验结果
研究问题
- RQ1在何种条件下,正则化的伪后验 $̄Pi_{γ}$ 在 $p \to \infty$ 且 $γ \downarrow 0$ 时收缩至真实参数?
- RQ2针对 $̄Pi_{γ}$ 的MCMC算法的混合时间如何依赖于设计矩阵的相干性?
- RQ3初始化对高维稀疏-密集模型中MCMC采样器收敛性有何影响?
- RQ4MCMC算法的计算复杂度如何随维度 $p$ 和稀疏度水平 $s_\star$ 变化?
- RQ5在初始化不佳或相干性较高的情况下,所提出的MCMC算法与现有Gibbs采样器在混合性能上如何比较?
主要发现
- 在充分条件下,即使在高维设置下,正则化的伪后验 $̄Pi_{γ}$ 仍以与真实后验相同的速率收缩至真实参数。
- 针对 $̄Pi_{γ}$ 的MCMC算法的混合时间对设计矩阵的相干性和链的初始化方式均极为敏感。
- 在有利情况下——低相干性和良好初始化——MCMC算法的计算复杂度与 $O(pe^{s_\star^2})$ 成正比,其中 $s_\star$ 为真实稀疏度水平。
- 模拟结果证实,当初始化不佳或设计矩阵相干性较高时,所提出的MCMC算法与Narisetti和He(2014)提出的Gibbs采样器均表现出较差的混合性能。
- $̄Pi_{γ}$ 的理论收缩速率与真实后验在线性回归模型中一致,验证了其作为计算可行替代方案的有效性。
- 本研究表明,设计矩阵的特性与初始化对高维稀疏-密集模型中后验的有效探索具有决定性影响。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。