Skip to main content
QUICK REVIEW

[Paper Review] Frequentist Consistency of Generalized Variational Inference

Jeremias Knoblauch|arXiv (Cornell University)|Dec 10, 2019
Gaussian Processes and Bayesian Inference50 references4 citations
TL;DR

This paper establishes frequentist consistency for generalized variational inference by proving that the approximate posterior distribution converges in distribution to a Dirac delta at the true parameter value. Under regularity conditions, the method ensures that the sequence of approximate posteriors becomes asymptotically concentrated around the true parameter, with error terms vanishing almost surely as sample size increases.

ABSTRACT

This paper investigates Frequentist consistency properties of the posterior distributions constructed via Generalized Variational Inference (GVI). A number of generic and novel strategies are given for proving consistency, relying on the theory of $Γ$-convergence. Specifically, this paper shows that under minimal regularity conditions, the sequence of GVI posteriors is consistent and collapses to a point mass at the population-optimal parameter value as the number of observations goes to infinity. The results extend to the latent variable case without additional assumptions and hold under misspecification. Lastly, the paper explains how to apply the results to a selection of GVI posteriors with especially popular variational families. For example, consistency is established for GVI methods using the mean field normal variational family, normal mixtures, Gaussian process variational families as well as neural networks indexing a normal (mixture) distribution.

Motivation & Objective

  • To establish frequentist consistency for generalized variational inference in Bayesian nonparametric and statistical learning settings.
  • To show that the sequence of approximate posteriors converges in distribution to a point mass at the true parameter value.
  • To derive sufficient conditions under which the approximation error terms vanish almost surely as the sample size grows.
  • To formalize the convergence of variational objectives using Γ-convergence and equi-coercivity.

Proposed method

  • Uses Γ-convergence of the variational objective functions $\overline{F}_n$ to $\mathbb{E}_q[\mathbb{E}_\mu[\ell(\bm{\theta}, \bm{x})]]$ under mild regularity assumptions.
  • Establishes equi-coercivity of $\overline{F}_n$ via Lemmas on $\Psi$ and $\overline{F}_n$ to ensure uniform lower bounds on the objective.
  • Defines $\varepsilon_n$-minimizers as sequences $q_n$ such that $\overline{F}_n(q_n) \leq \inf_q \overline{F}_n(q) + \varepsilon_n$, with $\varepsilon_n$ given by empirical deviation from population expectation.
  • Proves $\varepsilon_n \to 0$ almost surely under the data-generating measure $\mu$, using almost sure convergence of empirical averages.
  • Applies Corollary 7.24 from [gammaConvergence] to conclude $\overline{q}_n \overset{\mathcal{D}}{\to} \delta_{\bm{\theta}^*}$ almost surely.
  • Relies on Fubini-type arguments and almost sure finiteness of loss functions under the data measure $\mu$ to ensure well-definedness of expectations.

Experimental results

Research questions

  • RQ1Under what conditions does generalized variational inference produce posterior approximations that are consistent in the frequentist sense?
  • RQ2How can the convergence of the variational objective $\overline{F}_n$ be established using Γ-convergence and equi-coercivity?
  • RQ3What conditions ensure that the approximation error $\varepsilon_n$ vanishes almost surely as $n \to \infty$?
  • RQ4Can the sequence of approximate posteriors $\overline{q}_n$ be shown to converge in distribution to a Dirac delta at the true parameter $\bm{\theta}^*$?
  • RQ5How does the interplay between prior informativeness and empirical risk minimization affect the consistency of the variational posterior?

Key findings

  • The sequence of approximate posteriors $\overline{q}_n$ converges in distribution to a Dirac delta at the true parameter $\bm{\theta}^*$, i.e., $\overline{q}_n \overset{\mathcal{D}}{\to} \delta_{\bm{\theta}^*}$, under the stated assumptions.
  • The error term $\varepsilon_n = 2\left| \mathbb{E}_{\overline{q}_n}\left[\frac{1}{n}\sum_{i=1}^n \ell(\bm{\theta}, x_i) \right] - \mathbb{E}_\mu[\ell(\bm{\theta}, \bm{x})] \right|$ converges to zero almost surely as $n \to \infty$.
  • The variational objective $\overline{F}_n$ is equi-coercive, ensuring that minimizing sequences remain bounded in a suitable topology.
  • Γ-convergence of $\overline{F}_n$ to the population risk $\mathbb{E}_q[\mathbb{E}_\mu[\ell(\bm{\theta}, \bm{x})]]$ is established under Assumptions LABEL:AS:min_exists, LABEL:AS:dirac, LABEL:AS:D, and LABEL:AS:suitable.
  • The existence of a finite minimizer ensures that the prior is not infinitely bad and that posterior approximations improve upon the prior.
  • The convergence of $\overline{q}_n$ to $\delta_{\bm{\theta}^*}$ is proven using the combination of Γ-convergence, equi-coercivity, and almost sure convergence of empirical averages.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.