Skip to main content
QUICK REVIEW

[论文解读] Asymptotic Consistency of $α-$Rényi-Approximate Posteriors

Prateek Jaiswal, Vinayak Rao|arXiv (Cornell University)|Feb 5, 2019
advanced mathematical theories被引用 5
一句话总结

本文建立了 α-Rényi 近似后验的渐近一致性,这是一种变分贝叶斯方法,通过最小化 α-Rényi 散度(α > 1)来生成比传统基于 KL 散度的变分推断更宽泛、更鲁棒的后验近似。主要贡献在于识别出充分条件——核心在于存在一个具有适当收敛速率的‘良好’近似分布序列——在此条件下,随着数据量增加,这些近似能一致地恢复真实后验。

ABSTRACT

We study the asymptotic consistency properties of $α$-Rényi approximate posteriors, a class of variational Bayesian methods that approximate an intractable Bayesian posterior with a member of a tractable family of distributions, the member chosen to minimize the $α$-Rényi divergence from the true posterior. Unique to our work is that we consider settings with $α> 1$, resulting in approximations that upperbound the log-likelihood, and consequently have wider spread than traditional variational approaches that minimize the Kullback-Liebler (KL) divergence from the posterior. Our primary result identifies sufficient conditions under which consistency holds, centering around the existence of a 'good' sequence of distributions in the approximating family that possesses, among other properties, the right rate of convergence to a limit distribution. We further characterize the good sequence by demonstrating that a sequence of distributions that converges too quickly cannot be a good sequence. We also extend our analysis to the setting where $α$ equals one, corresponding to the minimizer of the reverse KL divergence, and to models with local latent variables. We also illustrate the existence of good sequence with a number of examples. Our results complement a growing body of work focused on the frequentist properties of variational Bayesian methods.

研究动机与目标

  • 建立理论条件,以确保 α-Rényi 近似后验在 α > 1 时具有渐近一致性。
  • 刻画确保一致性的‘良好’近似分布序列的性质,区分其与收敛过快的序列。
  • 将分析扩展至反向 KL 情况(α = 1)以及具有局部潜变量的模型。
  • 为在变分推断中使用 α-Rényi 散度提供理论依据,尤其适用于需要更宽后验分布的情形。
  • 通过在一般正则条件下建立一致性,补充现有对变分贝叶斯方法的频率学派分析。

提出的方法

  • 提出 α-Rényi 近似后验作为真实后验与可处理分布族中某成员之间 α-Rényi 散度的最小化器。
  • 使用 α > 1 的 α-Rényi 散度,生成上界对数似然的近似,因此其尾部比基于 KL 的方法更重。
  • 通过一个‘良好’的近似分布序列实现一致性,该序列以与模型正则性相容的速率收敛至真实后验,避免过快收敛。
  • 应用局部渐近正态性(LAN)条件,并利用 Bickel 和 Kleijn(2012)的结果,推导出真实模型与估计模型下似然比的渐近等价性。
  • 利用联合后验的结构并针对潜变量进行边际化,推导出 α-Rényi 散度的界,尤其适用于具有局部潜变量的模型。
  • 采用适配于 Rényi 散度的证据下界(ELBO)框架,表明最小化 α-Rényi 散度可导出边际似然的下界。

实验结果

研究问题

  • RQ1随着样本量增加,α-Rényi 近似后验在何种条件下收敛至真实后验?
  • RQ2‘良好’的近似分布序列与收敛过快而无法确保一致性的序列之间有何区别?
  • RQ3α-Rényi 变分推断的渐近行为与标准 KL 基变分推断(α = 1)相比如何?
  • RQ4一致性结果能否扩展至后验为高维的局部潜变量模型?
  • RQ5LAN 条件在建立真实模型与近似模型下似然比渐近等价性中起什么作用?

主要发现

  • 本文证明,若存在一个‘良好’的近似分布序列,其以与模型正则性相容的速率收敛至真实后验,则 α-Rényi 近似后验具有渐近一致性。
  • 收敛过快的序列不能被视为‘良好’序列,因其无法捕捉后验的完整范围,导致估计不一致。
  • 对于 α > 1,该方法生成的后验近似比基于 KL 的变分推断更宽泛,有利于捕捉方差和高阶矩中的不确定性。
  • 分析可扩展至反向 KL 情况(α = 1),在相同充分条件下仍保持一致性,从而统一了两种散度的处理。
  • 通过示例验证了在常见统计模型(包括具有局部潜变量的模型)中‘良好’序列的存在性。
  • 本文证明,在 s-LAN 条件下且近似序列适当收敛时,α-Rényi 散度以概率收敛至零,从而确保了一致性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。