[Paper Review] An approach to large-scale Quasi-Bayesian inference with spike-and-slab priors
This paper proposes a scalable quasi-Bayesian inference framework using spike-and-slab priors with a novel sparsification trick in the quasi-likelihood, enabling efficient MCMC and variational inference in high-dimensional settings. The key contribution is theoretical guarantees on posterior contraction and variational approximation accuracy, validated in Gaussian graphical models and sparse PCA.
We propose a general framework using spike-and-slab prior distributions to aid with the development of high-dimensional Bayesian inference. Our framework allows inference with a general quasi-likelihood function. We show that highly efficient and scalable Markov Chain Monte Carlo (MCMC) algorithms can be easily constructed to sample from the resulting quasi-posterior distributions. We study the large scale behavior of the resulting quasi-posterior distributions as the dimension of the parameter space grows, and we establish several convergence results. In large-scale applications where computational speed is important, variational approximation methods are often used to approximate posterior distributions. We show that the contraction behaviors of the quasi-posterior distributions can be exploited to provide theoretical guarantees for their variational approximations. We illustrate the theory with some simulation results from Gaussian graphical models, and sparse principal component analysis.
Motivation & Objective
- Address the computational intractability of point-mass spike-and-slab priors in high-dimensional Bayesian inference by replacing the point mass with a small-variance Gaussian.
- Develop a general framework for large-scale quasi-Bayesian inference using a flexible quasi-likelihood function, applicable beyond standard likelihoods.
- Establish theoretical convergence results for the resulting quasi-posterior distribution as dimensionality grows.
- Provide theoretical justification for variational approximations by exploiting posterior contraction behavior.
- Demonstrate scalability and accuracy through simulations in Gaussian graphical models and sparse principal component analysis.
Proposed method
- Use a spike-and-slab prior with Gaussian components: slab (precision ρ₁) for active variables (δⱼ = 1), spike (precision ρ₀ → 0) for inactive ones (δⱼ = 0), replacing point masses with small-variance Gaussians.
- Introduce a sparsified quasi-likelihood ℓ(θ₆; z) in the posterior, where θ₆ is the componentwise product of θ and δ, to improve statistical performance while preserving computational speed.
- Construct a scalable Gibbs sampler for posterior sampling by updating δ and θ conditionally, using normal and Bernoulli proposals with Metropolis-Hastings corrections.
- Develop a variational approximation (VA) algorithm using a mean-field template, with updates for α (inclusion probabilities), μ (mean), and C (covariance), leveraging a fixed support template δ^(i).
- Use the Bernstein-von Mises approximation to justify the asymptotic normality of the posterior and support variational approximation consistency.
- Leverage the sparsification trick to reduce computational cost in MCMC by focusing only on active variables in the likelihood evaluation.
Experimental results
Research questions
- RQ1Can a sparsified quasi-likelihood with spike-and-slab priors lead to faster and more scalable MCMC sampling in high-dimensional models?
- RQ2Under what conditions does the quasi-posterior concentrate around the true sparse parameter and its support?
- RQ3How can posterior contraction rates be used to justify the accuracy of variational approximations in high-dimensional settings?
- RQ4Does the proposed framework maintain optimal frequentist properties (e.g., variable selection consistency) while enabling scalable computation?
- RQ5How do the proposed MCMC and variational algorithms compare in terms of computational efficiency and estimation accuracy on real-world models like Gaussian graphical models and sparse PCA?
Key findings
- The sparsification trick in the quasi-likelihood significantly accelerates MCMC computation by reducing the number of active variables in likelihood evaluation.
- Theoretical results show that the quasi-posterior contracts around the true parameter (δ⋆, θ⋆) at a near-optimal rate under strong signal conditions.
- Posterior contraction rates are used to derive non-asymptotic bounds on the Kullback-Leibler divergence between the true posterior and its variational approximation.
- The variational approximation achieves consistent estimation of the support and parameter values, with convergence rates matching those of the true posterior under mild regularity conditions.
- Simulations in Gaussian graphical models and sparse PCA show that the proposed MCMC and variational algorithms achieve high accuracy and computational efficiency, even in high-dimensional regimes.
- The method achieves near-optimal variable selection performance, with the posterior mass concentrating on the true support δ⋆ as dimension p grows.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.