Skip to main content
QUICK REVIEW

[Paper Review] Convergence of Stein Variational Gradient Descent under a Weaker Smoothness Condition

Lukang Sun, Avetik Karagulyan|arXiv (Cornell University)|Jun 1, 2022
Markov Chains and Monte Carlo Methods66 references4 citations
TL;DR

This paper establishes convergence of Stein Variational Gradient Descent (SVGD) under a weaker smoothness condition—(L₀, L₁)-smoothness—instead of the standard L-smoothness assumption. By introducing trajectory-independent auxiliary conditions, the authors prove a descent lemma for the KL divergence and derive a complexity bound in terms of the Stein Fisher information, enabling convergence analysis for distributions with polynomial potentials, such as those with degree >2.

ABSTRACT

Stein Variational Gradient Descent (SVGD) is an important alternative to the Langevin-type algorithms for sampling from probability distributions of the form $π(x) \propto \exp(-V(x))$. In the existing theory of Langevin-type algorithms and SVGD, the potential function $V$ is often assumed to be $L$-smooth. However, this restrictive condition excludes a large class of potential functions such as polynomials of degree greater than $2$. Our paper studies the convergence of the SVGD algorithm for distributions with $(L_0,L_1)$-smooth potentials. This relaxed smoothness assumption was introduced by Zhang et al. [2019a] for the analysis of gradient clipping algorithms. With the help of trajectory-independent auxiliary conditions, we provide a descent lemma establishing that the algorithm decreases the $\mathrm{KL}$ divergence at each iteration and prove a complexity bound for SVGD in the population limit in terms of the Stein Fisher information.

Motivation & Objective

  • To extend the theoretical convergence analysis of SVGD beyond the restrictive L-smoothness assumption on the potential function V.
  • To address the limitation of existing SVGD theory, which excludes polynomial potentials of degree >2 due to the L-smoothness requirement.
  • To establish a descent lemma for the KL divergence in the population limit using trajectory-independent conditions.
  • To derive a complexity bound for SVGD in terms of the Stein Fisher information under the relaxed (L₀, L₁)-smoothness condition.

Proposed method

  • Introduce a weaker smoothness condition, (L₀, L₁)-smoothness, inspired by Zhang et al. (2019a), applicable to polynomial potentials.
  • Define a trajectory-independent auxiliary condition to avoid reliance on path-specific information in convergence analysis.
  • Establish a descent lemma showing that the KL divergence decreases at each SVGD iteration under the new smoothness condition.
  • Use the Stein Fisher information as a key quantity to derive a complexity bound for convergence in the population limit.
  • Leverage Grönwall’s inequality and weak convergence criteria involving distant dissipativity and inverse multiquadratic kernels to ensure convergence to the target measure.
  • Apply a variational formulation of SVGD as gradient descent on the KL divergence in the space of probability measures.

Experimental results

Research questions

  • RQ1Can SVGD be proven to converge under a weaker smoothness condition than L-smoothness?
  • RQ2Does the descent of the KL divergence in SVGD hold without relying on trajectory-dependent assumptions?
  • RQ3Can the complexity bound of SVGD be expressed in terms of the Stein Fisher information under (L₀, L₁)-smoothness?
  • RQ4Does the new framework allow convergence analysis for distributions with polynomial potentials of degree greater than 2?
  • RQ5What conditions ensure weak convergence of the SVGD iterates to the target distribution π under the relaxed smoothness assumption?

Key findings

  • The paper proves a descent lemma for SVGD in the population limit under (L₀, L₁)-smoothness, showing that the KL divergence decreases at each iteration.
  • The complexity bound for SVGD is derived in terms of the Stein Fisher information, providing a quantitative convergence rate under the new smoothness condition.
  • The analysis covers a broader class of distributions, including those with polynomial potentials of degree >2, which were excluded under the L-smoothness assumption.
  • The constant in the Wasserstein-KL bound depends on dimension, with λ_BV ≥ 2d^{1/p} for π(x) ∝ exp(−‖x‖^p), indicating potential dimension dependence.
  • Weak convergence of SVGD iterates to π is ensured under distant dissipativity of π and use of inverse multiquadratic kernels.
  • The framework removes the need for path-dependent trajectory conditions, improving the robustness and generality of convergence analysis.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.