[论文解读] Variational Bayes under Model Misspecification
本文在模型误设条件下建立了变分贝叶斯(VB)的渐近频率学派性质,证明VB后验及其均值会集中在使Kullback-Leibler散度最小化的参数值处。此外,研究进一步表明,在预测分布中,模型误设误差主导了变分近似误差,从而解释了尽管存在近似误差,VB在实践中仍表现出强大性能的原因。
Variational Bayes (VB) is a scalable alternative to Markov chain Monte Carlo (MCMC) for Bayesian posterior inference. Though popular, VB comes with few theoretical guarantees, most of which focus on well-specified models. However, models are rarely well-specified in practice. In this work, we study VB under model misspecification. We prove the VB posterior is asymptotically normal and centers at the value that minimizes the Kullback-Leibler (KL) divergence to the true data-generating distribution. Moreover, the VB posterior mean centers at the same value and is also asymptotically normal. These results generalize the variational Bernstein--von Mises theorem [29] to misspecified models. As a consequence of these results, we find that the model misspecification error dominates the variational approximation error in VB posterior predictive distributions. It explains the widely observed phenomenon that VB achieves comparable predictive accuracy with MCMC even though VB uses an approximating family. As illustrations, we study VB under three forms of model misspecification, ranging from model over-/under-dispersion to latent dimensionality misspecification. We conduct two simulation studies that demonstrate the theoretical results.
研究动机与目标
- 理解在实践中常见的模型误设情况下,变分贝叶斯(VB)的理论行为。
- 将变分贝叶斯-冯·米塞斯定理推广至误设模型,为VB后验及其均值提供渐近正态性结果。
- 量化模型误设与变分近似对VB预测误差的相对贡献。
- 通过在多种形式的模型误设下进行模拟,验证理论发现,包括过度/欠离散化和潜变量维度错误。
提出的方法
- 使用渐近频率学派工具,对误设条件下的VB进行理论分析,重点研究可分解密度的均场变分族。
- 推导VB后验与VB后验均值的极限分布,表明其依分布收敛至KL最小化点θ*的点质量分布。
- 利用Γ-收敛技术,建立在误设条件下变分目标泛函的收敛性。
- 通过局部拉普拉斯近似与集中不等式,推导出VB后验与VB后验均值在θ*附近的渐近正态性。
- 使用KL散度比较预测误差的分解,将其划分为模型误设与变分近似两部分。
- 通过Stan的HMC与VB实现进行模拟研究,验证在受控误设场景下的理论结果。
实验结果
研究问题
- RQ1在模型误设条件下,VB后验是否仍会集中在有意义的目标参数值处?
- RQ2VB后验均值是否渐近正态,且是否收敛至与VB后验相同的极限?
- RQ3模型误设误差与变分近似误差如何共同影响VB后验预测性能?
- RQ4变分贝叶斯-冯·米塞斯定理能否推广至误设模型?
- RQ5在不同形式的模型误设(如过度/欠离散化或潜变量维度错误)下,VB后验的极限行为如何?
主要发现
- VB后验渐近集中在θ*处,即最小化真实数据生成分布p₀(x)与模型之间KL散度的参数值。
- VB后验均值也以概率1收敛至θ*,且在该点附近渐近正态。
- 标准化后的VB后验√n(θ − θ*)依分布收敛至均值为零、协方差矩阵由模型Fisher信息量决定的正态分布。
- 在VB后预测分布中,模型误设导致的误差主导了变分近似带来的误差,从而解释了VB在预测中表现出的高精度。
- 通过两项模拟研究验证了理论结果,表明在包括过度/欠离散化和潜变量维度错误在内的各种误设类型下,均表现出一致的收敛性与预测性能。
- 建立了变分目标泛函的Γ-收敛性,证明在正则性条件下,极小化点收敛至KL最小化点θ*。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。