[Paper Review] Convergence of Stein Variational Gradient Descent under a Weaker Smoothness Condition
This paper establishes convergence of Stein Variational Gradient Descent (SVGD) under a weaker smoothness condition—(L₀, L₁)-smoothness—instead of the standard L-smoothness assumption. By introducing trajectory-independent auxiliary conditions, the authors prove a descent lemma for the KL divergence and derive a complexity bound in terms of the Stein Fisher information, enabling convergence analysis for distributions with polynomial potentials, such as those with degree >2.
Stein Variational Gradient Descent (SVGD) is an important alternative to the Langevin-type algorithms for sampling from probability distributions of the form $π(x) \propto \exp(-V(x))$. In the existing theory of Langevin-type algorithms and SVGD, the potential function $V$ is often assumed to be $L$-smooth. However, this restrictive condition excludes a large class of potential functions such as polynomials of degree greater than $2$. Our paper studies the convergence of the SVGD algorithm for distributions with $(L_0,L_1)$-smooth potentials. This relaxed smoothness assumption was introduced by Zhang et al. [2019a] for the analysis of gradient clipping algorithms. With the help of trajectory-independent auxiliary conditions, we provide a descent lemma establishing that the algorithm decreases the $\mathrm{KL}$ divergence at each iteration and prove a complexity bound for SVGD in the population limit in terms of the Stein Fisher information.
Motivation & Objective
- To extend the theoretical convergence analysis of SVGD beyond the restrictive L-smoothness assumption on the potential function V.
- To address the limitation of existing SVGD theory, which excludes polynomial potentials of degree >2 due to the L-smoothness requirement.
- To establish a descent lemma for the KL divergence in the population limit using trajectory-independent conditions.
- To derive a complexity bound for SVGD in terms of the Stein Fisher information under the relaxed (L₀, L₁)-smoothness condition.
Proposed method
- Introduce a weaker smoothness condition, (L₀, L₁)-smoothness, inspired by Zhang et al. (2019a), applicable to polynomial potentials.
- Define a trajectory-independent auxiliary condition to avoid reliance on path-specific information in convergence analysis.
- Establish a descent lemma showing that the KL divergence decreases at each SVGD iteration under the new smoothness condition.
- Use the Stein Fisher information as a key quantity to derive a complexity bound for convergence in the population limit.
- Leverage Grönwall’s inequality and weak convergence criteria involving distant dissipativity and inverse multiquadratic kernels to ensure convergence to the target measure.
- Apply a variational formulation of SVGD as gradient descent on the KL divergence in the space of probability measures.
Experimental results
Research questions
- RQ1Can SVGD be proven to converge under a weaker smoothness condition than L-smoothness?
- RQ2Does the descent of the KL divergence in SVGD hold without relying on trajectory-dependent assumptions?
- RQ3Can the complexity bound of SVGD be expressed in terms of the Stein Fisher information under (L₀, L₁)-smoothness?
- RQ4Does the new framework allow convergence analysis for distributions with polynomial potentials of degree greater than 2?
- RQ5What conditions ensure weak convergence of the SVGD iterates to the target distribution π under the relaxed smoothness assumption?
Key findings
- The paper proves a descent lemma for SVGD in the population limit under (L₀, L₁)-smoothness, showing that the KL divergence decreases at each iteration.
- The complexity bound for SVGD is derived in terms of the Stein Fisher information, providing a quantitative convergence rate under the new smoothness condition.
- The analysis covers a broader class of distributions, including those with polynomial potentials of degree >2, which were excluded under the L-smoothness assumption.
- The constant in the Wasserstein-KL bound depends on dimension, with λ_BV ≥ 2d^{1/p} for π(x) ∝ exp(−‖x‖^p), indicating potential dimension dependence.
- Weak convergence of SVGD iterates to π is ensured under distant dissipativity of π and use of inverse multiquadratic kernels.
- The framework removes the need for path-dependent trajectory conditions, improving the robustness and generality of convergence analysis.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.