Skip to main content
QUICK REVIEW

[论文解读] Primal Dual Interpretation of the Proximal Stochastic Gradient Langevin Algorithm

Adil Salim, Peter Richtárik|arXiv (Cornell University)|Jun 16, 2020
Markov Chains and Monte Carlo Methods参考文献 27被引用 6
一句话总结

本论文为从具有复合势函数的对数凹分布中进行采样,提供了近端随机梯度朗之万算法(PSGLA)的对偶-对偶解释,其中势函数为光滑凸函数与非光滑凸函数(可能取无穷大)之和。通过在 Wasserstein 空间中利用强对偶性,该工作在强凸性假设下建立了 2-Wasserstein 距离下的 ${\mathcal{O}}(1/\varepsilon^{2})$ 复杂度界,显著优于投影朗之万算法的 ${\mathcal{O}}(1/\varepsilon^{12})$ 界。

ABSTRACT

We consider the task of sampling with respect to a log concave probability distribution. The potential of the target distribution is assumed to be composite, extit{i.e.}, written as the sum of a smooth convex term, and a nonsmooth convex term possibly taking infinite values. The target distribution can be seen as a minimizer of the Kullback-Leibler divergence defined on the Wasserstein space ( extit{i.e.}, the space of probability measures). In the first part of this paper, we establish a strong duality result for this minimization problem. In the second part of this paper, we use the duality gap arising from the first part to study the complexity of the Proximal Stochastic Gradient Langevin Algorithm (PSGLA), which can be seen as a generalization of the Projected Langevin Algorithm. Our approach relies on viewing PSGLA as a primal dual algorithm and covers many cases where the target distribution is not fully supported. In particular, we show that if the potential is strongly convex, the complexity of PSGLA is $O(1/\varepsilon^2)$ in terms of the 2-Wasserstein distance. In contrast, the complexity of the Projected Langevin Algorithm is $O(1/\varepsilon^{12})$ in terms of total variation when the potential is convex.

研究动机与目标

  • 为分析近端随机梯度朗之万算法(PSGLA)在从具有复合势函数的对数凹分布中采样提供一个对偶-对偶框架。
  • 在 Wasserstein 空间中建立最小化 Kullback-Leibler 散度的强对偶性,从而支持 PSGLA 的复杂度分析。
  • 推导 PSGLA 在 2-Wasserstein 距离下的非渐近收敛速率,尤其针对目标分布不完全支持的情况。
  • 将复杂度界扩展至势函数 $G$ 可能取无穷大时的情形,涵盖约束采样问题。

提出的方法

  • 将采样问题表述为在 Wasserstein 空间中最小化 KL 散度,采用复合势 $V = F + G$,其中 $F$ 为光滑函数,$G$ 为非光滑凸函数。
  • 为 KL 最小化问题建立强对偶性,引入包含对偶变量和 $G$ 共轭函数的拉格朗日函数。
  • 将 PSGLA 视为一个对偶-对偶算法,其中原始迭代 $x^{k+1}$ 通过近端步长更新,对偶迭代 $y^{k+1}$ 由对偶问题导出。
  • 推导一个涉及当前分布与目标分布之间 2-Wasserstein 距离的递推不等式,包含对偶间隙项和噪声方差。
  • 利用 $F$ 的强凸性和 $G^*$ 的光滑性,通过对偶间隙控制 Wasserstein 距离的衰减。
  • 应用对随机梯度和噪声的浓度与矩不等式,以控制递推式中的残差项。

实验结果

研究问题

  • RQ1PSGLA 能否在复合对数凹分布采样的背景下被解释为对偶-对偶算法?
  • RQ2当 $G$ 为非光滑或取无穷大值时,PSGLA 在 2-Wasserstein 距离下的非渐近收敛速率为何?
  • RQ3在总变差距离和 Wasserstein 距离下,PSGLA 的复杂度与投影朗之万算法相比如何?
  • RQ4能否在 Wasserstein 空间中利用强对偶性,为具有约束支撑的采样算法推导更紧致的收敛界?

主要发现

  • 当势函数 $F$ 为强凸函数且 $G$ 为 $1/\lambda_{G^*}$-光滑时,PSGLA 在 2-Wasserstein 距离下达到 ${\mathcal{O}}(1/\varepsilon^{2})$ 的复杂度界。
  • 该复杂度显著优于当 $F$ 为凸且光滑时,投影朗之万算法在总变差距离下的 ${\mathcal{O}}(1/\varepsilon^{12})$ 界。
  • 由强对偶性结果引出的对偶间隙被用于推导控制 Wasserstein 距离衰减的递推不等式。
  • 即使对某些 $x$ 有 $G(x) = +\infty$,该分析依然成立,覆盖了目标分布支持于凸体的情形。
  • 该方法可扩展至带额外正则化项的随机近端朗之万算法(SPLA),在适当假设下保持相似的收敛速率。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。