Skip to main content
QUICK REVIEW

[Paper Review] Convergence Rates of Variational Posterior Distributions

Fengshuo Zhang, Chao Gao|arXiv (Cornell University)|Dec 7, 2017
Statistical Methods and Inference40 references10 citations
TL;DR

This paper establishes convergence rates for variational posterior distributions in nonparametric and high-dimensional models by deriving general conditions on the prior, likelihood, and variational class. It shows that the rate is the sum of the true posterior's convergence rate and a variational approximation error term, and under a novel prior mass condition for mixture-of-product priors, the approximation error is dominated by the statistical rate, ensuring optimal performance of mean-field variational inference.

ABSTRACT

We study convergence rates of variational posterior distributions for nonparametric and high-dimensional inference. We formulate general conditions on prior, likelihood, and variational class that characterize the convergence rates. Under similar "prior mass and testing" conditions considered in the literature, the rate is found to be the sum of two terms. The first term stands for the convergence rate of the true posterior distribution, and the second term is contributed by the variational approximation error. For a class of priors that admit the structure of a mixture of product measures, we propose a novel prior mass condition, under which the variational approximation error of the mean-field class is dominated by convergence rate of the true posterior. We demonstrate the applicability of our general results for various models, prior distributions and variational classes by deriving convergence rates of the corresponding variational posteriors.

Motivation & Objective

  • To establish theoretical convergence rates for variational posterior distributions in nonparametric and high-dimensional inference problems.
  • To identify general conditions on the prior, likelihood, and variational class that ensure convergence of the variational posterior to the true data-generating distribution.
  • To show that under a novel prior mass condition tailored to mixture-of-product priors, the variational approximation error is dominated by the statistical error of the true posterior.
  • To demonstrate the applicability of the framework across models including density estimation, Gaussian sequence models, and piecewise constant models.
  • To reveal a theoretical connection between empirical Bayes and variational Bayes via mean-field approximation.

Proposed method

  • The paper formulates a general convergence rate bound for variational posteriors as $ \epsilon_n^2 + \frac{1}{n} \inf_{Q \in \mathcal{S}} P_0^{(n)} D(Q \| \Pi(\cdot|X^{(n)})) $, where $ \epsilon_n^2 $ is the true posterior rate and the second term is the variational approximation error.
  • It introduces a new prior mass condition for mixture-of-product priors: $ \Pi(\otimes_j \widetilde{\Theta}_j) \geq \exp(-C_2 n \epsilon_n^2) $, under which the variational error is dominated by the true posterior rate.
  • The analysis leverages the 'prior mass and testing' framework from posterior contraction theory, adapting it to variational inference by preserving the same three core conditions: prior mass, testing, and support.
  • For the mean-field variational class $ \mathcal{S}_{\rm MF} $, the paper shows that the approximation error is controlled when the prior satisfies the new condition, ensuring the variational posterior achieves the same rate as the true posterior.
  • The paper establishes a connection between empirical Bayes and variational Bayes by showing that empirical Bayes procedures are equivalent to variational Bayes using a specially designed variational class.
  • It applies the framework to specific models, including the Gaussian sequence model and piecewise constant models, deriving explicit convergence rates under the proposed conditions.

Experimental results

Research questions

  • RQ1Under what conditions does the variational posterior converge to the true parameter at a rate comparable to the true posterior?
  • RQ2Can the variational approximation error be made negligible relative to the statistical error of the true posterior?
  • RQ3How should priors be constructed to ensure that mean-field variational inference achieves optimal convergence rates?
  • RQ4What is the theoretical relationship between empirical Bayes and variational Bayes procedures?
  • RQ5Can the 'prior mass and testing' framework be extended to analyze variational posteriors with minimal modification?

Key findings

  • The convergence rate of the variational posterior is bounded by $ \epsilon_n^2 + \frac{1}{n} \inf_{Q \in \mathcal{S}} P_0^{(n)} D(Q \| \Pi(\cdot|X^{(n)})) $, where the first term is the true posterior rate and the second is the variational approximation error.
  • For mixture-of-product priors and the mean-field variational class, the approximation error is dominated by the true posterior rate if the prior satisfies $ \Pi(\otimes_j \widetilde{\Theta}_j) \geq \exp(-C_2 n \epsilon_n^2) $.
  • In the Gaussian sequence model with a sparsity-inducing prior, the variational posterior achieves a rate of $ \min_k \left\{ \|\theta^* - \theta_0\|^2 + k \log n \right\} $, matching the optimal posterior rate.
  • The paper shows that empirical Bayes procedures are equivalent to variational Bayes using a specific variational class, suggesting shared theoretical properties.
  • The condition $ \Pi(\otimes_j \widetilde{\Theta}_j) \geq \exp(-C_2 n \epsilon_n^2) $ is both necessary and sufficient for the variational error to be negligible, and it is easy to verify in practice.
  • The framework applies broadly, including to density estimation and piecewise constant models, where the variational posterior achieves the minimax-optimal rate under the new prior mass condition.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.