Skip to main content
QUICK REVIEW

[论文解读] Challenges and Opportunities in High-dimensional Variational Inference

Akash Kumar Dhaka, Alejandro Catalina|OpenBU (Boston University)|Mar 1, 2021
Gaussian Processes and Bayesian Inference参考文献 28被引用 4
一句话总结

本文提出了一套框架,通过分析目标分布与变分分布之间密度比的预渐近尾部行为,利用重要性采样工具改进高维变分推断。研究发现,使用灵活的变分族(尤其是归一化流)时,仅使用KL散度可获得最可靠且精确的近似结果;而质量覆盖型散度(如包含式KL)由于不稳定性和收敛性差,在高维情况下表现不佳。

ABSTRACT

Current black-box variational inference (BBVI) methods require the user to make numerous design choices -- such as the selection of variational objective and approximating family -- yet there is little principled guidance on how to do so. We develop a conceptual framework and set of experimental tools to understand the effects of these choices, which we leverage to propose best practices for maximizing posterior approximation accuracy. Our approach is based on studying the pre-asymptotic tail behavior of the density ratios between the joint distribution and the variational approximation, then exploiting insights and tools from the importance sampling literature. Our framework and supporting experiments help to distinguish between the behavior of BBVI methods for approximating low-dimensional versus moderate-to-high-dimensional posteriors. In the latter case, we show that mass-covering variational objectives are difficult to optimize and do not improve accuracy, but flexible variational families can improve accuracy and the effectiveness of importance sampling -- at the cost of additional optimization challenges. Therefore, for moderate-to-high-dimensional posteriors we recommend using the (mode-seeking) exclusive KL divergence since it is the easiest to optimize, and improving the variational family or using model parameter transformations to make the posterior and optimal variational approximation more similar. On the other hand, in low-dimensional settings, we show that heavy-tailed variational families and mass-covering divergences are effective and can increase the chances that the approximation can be improved by importance sampling.

研究动机与目标

  • 解决在高维贝叶斯后验近似中,选择变分目标函数与近似族时缺乏系统性指导的问题。
  • 探究为何关于变分推断的低维直观认识在高维设置下会失效。
  • 基于预渐近收敛行为,构建概念性与实证性框架,用于评估黑箱变分推断(BBVI)方法的可靠性。
  • 为实践者提供在不同后验维度范围内优化变分推断的可操作建议。
  • 评估重要性采样修正方法(如PSIS)在不同条件下提升近似精度的有效性。

提出的方法

  • 分析真实后验与变分近似之间密度比 $ w(\theta) = p(\theta,Y)/q(\theta) $ 的预渐近尾部行为。
  • 利用重要性采样中的Pareto $ k $ 诊断,评估BBVI中梯度与散度估计器的可靠性。
  • 通过实证方法评估多种散度(排除式/包含式KL、$ \chi^2 $、$ \alpha $-散度、尾部自适应 $ f $-散度)与变分族(正态分布、学生t分布、归一化流)的表现。
  • 使用posteralldb中的模拟数据与真实数据集,比较不同方法在后验矩估计与预测似然方面的表现。
  • 应用PSIS(Pareto平滑重要性采样)对后验期望进行修正,并评估其在精度上的改进效果。
  • 采用模型重参数化方法,使后验结构与变分族对齐,从而提升近似质量。

实验结果

研究问题

  • RQ1为何质量覆盖型变分目标函数(如包含式KL)在高维后验中尽管理论上有吸引力,却无法提升精度?
  • RQ2密度比的预渐近行为如何影响黑箱变分推断中梯度估计器的可靠性?
  • RQ3在不同后验维度下,不同散度(如排除式KL与包含式KL)的相对优缺点是什么?
  • RQ4在 $ \hat{k} $ 诊断显示不稳定时,重要性采样修正方法(如PSIS)能在多大程度上提升变分近似的精度?
  • RQ5归一化流与重参数化策略在多大程度上影响高维后验近似的稳定性与精度?

主要发现

  • 在高维设置中,排除式KL散度由于优化稳定性更好且梯度估计方差更低,始终优于包含式KL及其他质量覆盖型散度。
  • 对于高维后验($ D > 10 $),结合排除式KL与PSIS修正的归一化流方法,在多种数据集上均能提供最精确的后验近似。
  • 在低维设置中($ D \leq 10 $),重尾变分族(如学生t分布)与质量覆盖型散度可提升近似精度,并增强重要性采样效果。
  • Pareto $ \hat{k} $ 诊断能可靠识别不稳定估计器:高 $ \hat{k} $ 值表明收敛性差,且原始估计与PSIS修正估计的可靠性均降低。
  • PSIS修正显著降低了均值与协方差估计的相对误差,尤其在归一化流近似中效果明显,证明其在 $ \hat{k} $ 较大时仍具重要价值。
  • 通过重参数化使后验结构与变分族对齐(如对漏斗形后验),可显著提升近似精度,即使在中等维度下亦然。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。