[Paper Review] Challenges and Opportunities in High-dimensional Variational Inference
This paper proposes a framework for improving high-dimensional variational inference by analyzing the pre-asymptotic tail behavior of density ratios between target and variational distributions, using tools from importance sampling. It finds that exclusive KL divergence with flexible variational families—especially normalizing flows—yields the most reliable and accurate approximations in high dimensions, while mass-covering divergences like inclusive KL fail due to instability and poor convergence.
Current black-box variational inference (BBVI) methods require the user to make numerous design choices -- such as the selection of variational objective and approximating family -- yet there is little principled guidance on how to do so. We develop a conceptual framework and set of experimental tools to understand the effects of these choices, which we leverage to propose best practices for maximizing posterior approximation accuracy. Our approach is based on studying the pre-asymptotic tail behavior of the density ratios between the joint distribution and the variational approximation, then exploiting insights and tools from the importance sampling literature. Our framework and supporting experiments help to distinguish between the behavior of BBVI methods for approximating low-dimensional versus moderate-to-high-dimensional posteriors. In the latter case, we show that mass-covering variational objectives are difficult to optimize and do not improve accuracy, but flexible variational families can improve accuracy and the effectiveness of importance sampling -- at the cost of additional optimization challenges. Therefore, for moderate-to-high-dimensional posteriors we recommend using the (mode-seeking) exclusive KL divergence since it is the easiest to optimize, and improving the variational family or using model parameter transformations to make the posterior and optimal variational approximation more similar. On the other hand, in low-dimensional settings, we show that heavy-tailed variational families and mass-covering divergences are effective and can increase the chances that the approximation can be improved by importance sampling.
Motivation & Objective
- To address the lack of principled guidance in selecting variational objectives and approximating families for high-dimensional Bayesian posterior approximation.
- To investigate why common low-dimensional intuitions about variational inference fail in high-dimensional settings.
- To develop a conceptual and empirical framework for evaluating the reliability of black-box variational inference (BBVI) methods based on pre-asymptotic convergence behavior.
- To provide actionable recommendations for practitioners on optimizing variational inference in diverse posterior dimensionality regimes.
- To assess the effectiveness of importance sampling corrections (e.g., PSIS) in improving approximation accuracy under varying conditions.
Proposed method
- Analyzes the pre-asymptotic tail behavior of the density ratio $ w(\theta) = p(\theta,Y)/q(\theta) $ between the true posterior and variational approximation.
- Leverages the Pareto $ k $ diagnostic from importance sampling to assess the reliability of gradient and divergence estimators in BBVI.
- Empirically evaluates multiple divergences (exclusive/inclusive KL, $ \chi^2 $, $ \alpha $-divergence, tail-adaptive $ f $-divergence) and variational families (Gaussian, Student-$ t $, normalizing flows).
- Uses simulated and real-world datasets from posteriordb to compare posterior moment estimation and predictive likelihood across methods.
- Applies PSIS (Pareto-smoothed importance sampling) to correct posterior expectations and assess improvement in accuracy.
- Employs model reparameterization to align the posterior structure with the variational family, improving approximation quality.
Experimental results
Research questions
- RQ1Why do mass-covering variational objectives like inclusive KL fail to improve accuracy in high-dimensional posteriors despite theoretical appeal?
- RQ2How does the pre-asymptotic behavior of density ratios affect the reliability of gradient estimators in black-box variational inference?
- RQ3What are the relative strengths and weaknesses of different divergences (e.g., exclusive vs. inclusive KL) across varying posterior dimensionality?
- RQ4To what extent can importance sampling corrections (e.g., PSIS) improve the accuracy of variational approximations, especially when $ \hat{k} $ diagnostics indicate instability?
- RQ5How do normalizing flows and reparameterization strategies affect the stability and accuracy of high-dimensional posterior approximations?
Key findings
- Exclusive KL divergence consistently outperforms inclusive KL and other mass-covering divergences in high-dimensional settings due to better optimization stability and lower variance in gradient estimators.
- For high-dimensional posteriors ($ D > 10 $), normalizing flows combined with exclusive KL and PSIS correction yield the most accurate posterior approximations across diverse datasets.
- In low-dimensional settings ($ D \leq 10 $), heavy-tailed variational families (e.g., Student-$ t $) and mass-covering divergences can improve approximation accuracy and enhance the effectiveness of importance sampling.
- The Pareto $ \hat{k} $ diagnostic reliably identifies unstable estimators: high $ \hat{k} $ values indicate poor convergence and reduced reliability of both raw and PSIS-corrected estimates.
- PSIS correction significantly reduces relative error in mean and covariance estimates—especially for normalizing flow approximations—demonstrating its value even when $ \hat{k} $ is large.
- Model reparameterization that aligns the posterior structure with the variational family (e.g., for funnel-shaped posteriors) can dramatically improve approximation accuracy, even in moderate dimensions.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.