Skip to main content
QUICK REVIEW

[论文解读] Asymptotic Consistency of $\alpha$-{R}\'enyi-Approximate Posteriors

Prateek Jaiswal, Vinayak Rao|arXiv (Cornell University)|Feb 5, 2019
Gaussian Processes and Bayesian Inference参考文献 27被引用 5
一句话总结

本文建立了 α-Rényi 近似后验的渐近一致性,这是一种变分贝叶斯方法,通过最小化 α-Rényi 散度(α > 1)来生成比传统基于 KL 散度的变分推断更宽泛、更鲁棒的后验近似。关键贡献在于识别出充分条件——核心在于存在一个具有适当收敛速率的‘良好’近似分布序列——在此条件下,随着数据量增加,α-Rényi 近似后验能一致地收敛于真实后验。

ABSTRACT

We study the asymptotic consistency properties of $\\alpha$-R\\'enyi approximate posteriors, a class of variational Bayesian methods that approximate an intractable Bayesian posterior with a member of a tractable family of distributions, the member chosen to minimize the $\\alpha$-R\\'enyi divergence from the true posterior. Unique to our work is that we consider settings with $\\alpha > 1$, resulting in approximations that upperbound the log-likelihood, and consequently have wider spread than traditional variational approaches that minimize the Kullback-Liebler (KL) divergence from the posterior. Our primary result identifies sufficient conditions under which consistency holds, centering around the existence of a 'good' sequence of distributions in the approximating family that possesses, among other properties, the right rate of convergence to a limit distribution. We further characterize the good sequence by demonstrating that a sequence of distributions that converges too quickly cannot be a good sequence. We also extend our analysis to the setting where $\\alpha$ equals one, corresponding to the minimizer of the reverse KL divergence, and to models with local latent variables. We also illustrate the existence of good sequence with a number of examples. Our results complement a growing body of work focused on the frequentist properties of variational Bayesian methods.

研究动机与目标

  • 建立理论条件,以确保 α-Rényi 近似后验在 α > 1 时具有渐近一致性。
  • 通过使用一种鼓励更宽泛、更保守近似的散度度量,解决传统变分推断低估后验方差的局限性。
  • 刻画在 α-Rényi 散度最小化背景下,‘良好’近似分布序列所需满足的必要属性,包括其收敛速率,以确保一致性。
  • 将分析扩展至反向 KL 情况(α = 1)以及具有局部潜变量的模型。
  • 为 α-Rényi 变分推断的频率学性质提供理论依据,补充该领域现有研究。

提出的方法

  • 本文分析了 α-Rényi 散度最小化作为变分推断方法,其中近似后验被选择为最小化 Dα(π(θ|Xn) || q(θ))(α > 1)。
  • 引入了‘良好’近似分布序列 {q̄n(θ)} 的概念,其以与模型正则性条件相容的速率收敛于真实后验。
  • 理论分析依赖于局部渐近正态性(LAN)条件,并使用霍尔丁格型度量来控制近似族的收敛性。
  • 证明了收敛过快的序列无法构成‘良好’序列,从而为一致性确立了收敛速率的必要下界。
  • 通过在分析中将潜变量固定为其真实值,将该框架扩展至具有局部潜变量的模型。
  • 证明技术利用了 α-Rényi 散度上界于对数似然的事实,从而导致比 KL 最小化更重尾的近似。

实验结果

研究问题

  • RQ1在何种条件下,α > 1 的 α-Rényi 变分推断能产生渐近一致的后验近似?
  • RQ2在 α-Rényi 散度最小化背景下,近似分布序列必须满足哪些属性才能被视为‘良好’?
  • RQ3α-Rényi 近似理论一致性是否可扩展至具有局部潜变量的模型?
  • RQ4近似序列的收敛速率如何影响 α-Rényi 后验的一致性?
  • RQ5在渐近行为方面,α-Rényi 散度与反向 KL 散度(α = 1)之间存在何种关系?

主要发现

  • 本文证明,若存在一个‘良好’的近似分布序列,其以适当的速率收敛于真实后验,则 α-Rényi 近似后验的渐近一致性成立。
  • 收敛过快的序列无法构成‘良好’序列,意味着一致性要求收敛速率存在必要下界。
  • 对于 α > 1,α-Rényi 散度产生的后验近似比 KL 最小化更宽泛,因其上界于对数似然并惩罚覆盖不足。
  • 分析已扩展至反向 KL 情况(α = 1),表明在类似条件下仍具有一致性,从而统一了对两种散度的处理。
  • 通过示例验证了标准模型中‘良好’序列的存在性,支持了理论框架。
  • 本文确认,在包括 LAN 和霍尔丁格收敛在内的正则性条件下,当 n → ∞ 时,α-Rényi 近似后验依分布收敛于真实后验。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。