Skip to main content
QUICK REVIEW

[论文解读] Frequentist Consistency of Generalized Variational Inference

Jeremias Knoblauch|arXiv (Cornell University)|Dec 10, 2019
Gaussian Processes and Bayesian Inference参考文献 50被引用 4
一句话总结

本文通过证明近似后验分布以概率收敛于真实参数值处的狄拉克测度,建立了广义变分推断的频率学一致性。在正则性条件下,该方法确保近似后验序列渐近集中于真实参数,且随着样本量增加,误差项几乎必然趋于零。

ABSTRACT

This paper investigates Frequentist consistency properties of the posterior distributions constructed via Generalized Variational Inference (GVI). A number of generic and novel strategies are given for proving consistency, relying on the theory of $Γ$-convergence. Specifically, this paper shows that under minimal regularity conditions, the sequence of GVI posteriors is consistent and collapses to a point mass at the population-optimal parameter value as the number of observations goes to infinity. The results extend to the latent variable case without additional assumptions and hold under misspecification. Lastly, the paper explains how to apply the results to a selection of GVI posteriors with especially popular variational families. For example, consistency is established for GVI methods using the mean field normal variational family, normal mixtures, Gaussian process variational families as well as neural networks indexing a normal (mixture) distribution.

研究动机与目标

  • 在贝叶斯非参数和统计学习设置中,建立广义变分推断的频率学一致性。
  • 证明近似后验序列以分布收敛于真实参数值处的点质量。
  • 推导近似误差项在样本量增长时以几乎必然方式趋于零的充分条件。
  • 通过Γ-收敛和等度强制性形式化变分目标的收敛性。

提出的方法

  • 在温和正则性假设下,证明变分目标函数 $\overline{F}_n$ 的Γ-收敛性,其极限为 $\mathbb{E}_q[\mathbb{E}_\mu[\ell(\bm{\theta}, \bm{x})]]$。
  • 通过关于 $\Psi$ 和 $\overline{F}_n$ 的引理,建立 $\overline{F}_n$ 的等度强制性,以确保目标函数的统一下界。
  • 将 $\varepsilon_n$-极小化器定义为满足 $\overline{F}_n(q_n) \leq \inf_q \overline{F}_n(q) + \varepsilon_n$ 的序列 $q_n$,其中 $\varepsilon_n$ 由经验偏差与总体期望的差值给出。
  • 利用经验平均的几乎必然收敛性,证明 $\varepsilon_n \to 0$ 几乎必然成立(在数据生成测度 $\mu$ 下)。
  • 应用文献 [gammaConvergence] 中的推论 7.24,得出 $\overline{q}_n \overset{\mathcal{D}}{\to} \delta_{\bm{\theta}^*}$ 几乎必然成立。
  • 依赖Fubini型论证以及在数据测度 $\mu$ 下损失函数的几乎必然有限性,以确保期望的良定义性。

实验结果

研究问题

  • RQ1在何种条件下,广义变分推断能产生在频率学意义上一致的后验近似?
  • RQ2如何通过Γ-收敛性和等度强制性建立变分目标 $\overline{F}_n$ 的收敛性?
  • RQ3在什么条件下,近似误差 $\varepsilon_n$ 能够在 $n \to \infty$ 时以几乎必然方式趋于零?
  • RQ4能否证明近似后验序列 $\overline{q}_n$ 以分布收敛于真实参数 $\bm{\theta}^*$ 处的狄拉克测度?
  • RQ5先验信息的丰富程度与经验风险最小化之间的相互作用,如何影响变分后验的一致性?

主要发现

  • 在给定假设下,近似后验序列 $\overline{q}_n$ 以分布收敛于真实参数 $\bm{\theta}^*$ 处的狄拉克测度,即 $\overline{q}_n \overset{\mathcal{D}}{\to} \delta_{\bm{\theta}^*}$。
  • 误差项 $\varepsilon_n = 2\left| \mathbb{E}_{\overline{q}_n}\left[\frac{1}{n}\sum_{i=1}^n \ell(\bm{\theta}, x_i) \right] - \mathbb{E}_\mu[\ell(\bm{\theta}, \bm{x})] \right|$ 几乎必然收敛于零,当 $n \to \infty$ 时。
  • 变分目标 $\overline{F}_n$ 具有等度强制性,确保最小化序列在适当拓扑下有界。
  • 在假设 LABEL:AS:min_exists、LABEL:AS:dirac、LABEL:AS:D 和 LABEL:AS:suitable 下,建立了 $\overline{F}_n$ 对总体风险 $\mathbb{E}_q[\mathbb{E}_\mu[\ell(\bm{\theta}, \bm{x})]]$ 的Γ-收敛性。
  • 存在有限最小化器,确保先验不会无限差,并且后验近似优于先验。
  • 通过Γ-收敛性、等度强制性和经验平均的几乎必然收敛性,证明了 $\overline{q}_n$ 收敛于 $\delta_{\bm{\theta}^*}$。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。