Skip to main content
QUICK REVIEW

[论文解读] On the Convergence of Langevin Monte Carlo: The Interplay between Tail Growth and Smoothness

Murat A. Erdogdu, Rasa Hosseinzadeh|arXiv (Cornell University)|May 27, 2020
Markov Chains and Monte Carlo Methods参考文献 58被引用 20
一句话总结

本论文在弱光滑性与尾部增长条件下,建立了非调整Langevin蒙特卡洛(LMC)算法的收敛速率:对于具有$α$阶尾部增长($\alpha \in [1,2]$)且梯度为$\beta$-Hölder连续的势函数,LMC在KL散度上达到$\epsilon$-精度所需的步数为$\widetilde{\mathcal{O}}\big(d^{\frac{1}{\beta} + \frac{1+\beta}{\beta}(\frac{2}{\alpha} - \mathbbm{1}_{\{\alpha \neq 1\}})} \epsilon^{-\frac{1}{\beta}}\big)$。关键洞见在于,当$\alpha \geq 1$时,$\epsilon$-依赖性仅取决于光滑性$\beta$,而与尾部增长$\alpha$无关,即使在非凸势函数下,该结果也恢复了梯度Lipschitz连续($\beta=1$)时的最佳已知速率。

ABSTRACT

We study sampling from a target distribution ${ν_* = e^{-f}}$ using the unadjusted Langevin Monte Carlo (LMC) algorithm. For any potential function $f$ whose tails behave like ${\|x\|^α}$ for ${α\in [1,2]}$, and has $β$-Hölder continuous gradient, we prove that ${\widetilde{\mathcal{O}} \Big(d^{\frac{1}β+\frac{1+β}β(\frac{2}α - \boldsymbol{1}_{\{α eq 1\}})} ε^{-\frac{1}β}\Big)}$ steps are sufficient to reach the $ε$-neighborhood of a $d$-dimensional target distribution $ν_*$ in KL-divergence. This convergence rate, in terms of $ε$ dependency, is not directly influenced by the tail growth rate $α$ of the potential function as long as its growth is at least linear, and it only relies on the order of smoothness $β$. One notable consequence of this result is that for potentials with Lipschitz gradient, i.e. $β=1$, our rate recovers the best known rate ${\widetilde{\mathcal{O}}(dε^{-1})}$ which was established for strongly convex potentials in terms of $ε$ dependency, but we show that the same rate is achievable for a wider class of potentials that are degenerately convex at infinity. The growth rate $α$ starts to have an effect on the established rate in high dimensions where $d$ is large; furthermore, it recovers the best-known dimension dependency when the tail growth of the potential is quadratic, i.e. ${α= 2}$, in the current setup. Our framework allows for finite perturbations, and any order of smoothness ${β\in(0,1]}$; consequently, our results are applicable to a wide class of non-convex potentials that are weakly smooth and exhibit at least linear tail growth.

研究动机与目标

  • 在势函数的结构假设最小化的情况下,建立非调整Langevin蒙特卡洛(LMC)的收敛速率。
  • 分析尾部增长速率$\alpha$与梯度光滑性$\beta$如何共同影响高维采样中LMC的收敛性。
  • 将已知的收敛保证从强凸或光滑势函数扩展到更广泛的非凸、弱光滑分布类,其尾部增长至少为线性。
  • 提出一种依赖于矩的修正对数 Sobolev 不等式(MLSI)框架,以捕捉非对数凹目标下尾部行为与光滑性之间的相互作用。
  • 证明当$\beta=1$时,$\widetilde{\mathcal{O}}(d\epsilon^{-1})$的速率对无穷远处退化凸势函数也成立,而不仅限于强凸情况。

提出的方法

  • 针对在无穷远处具有凸性且具有$\alpha$-增长尾部的目标分布,证明具有显式常数的依赖于矩的修正对数 Sobolev 不等式。
  • 在$\alpha$-发散性条件下,建立LMC马氏链的线性发散矩估计,将矩的增长与收敛速率联系起来。
  • 利用从矩估计中导出的微分不等式,并通过迭代控制LMC分布与目标分布之间的KL散度。
  • 调节矩的阶数以平衡矩增长与收敛性之间的权衡,从而实现对KL误差的紧密控制。
  • 应用Holley-Stroock扰动引理,将结果推广至凸势函数的有限扰动,从而扩大在非凸目标上的适用性。
  • 利用Talagrand不等式与相对Fisher信息,将KL散度与Wasserstein距离关联,实现多度量下的收敛性分析。

实验结果

研究问题

  • RQ1势函数的尾部增长速率$\alpha$在高维下如何影响非调整LMC的收敛速率?
  • RQ2对于梯度Lipschitz连续($\beta=1$)的情况,$\widetilde{\mathcal{O}}(d\epsilon^{-1})$的收敛速率是否可推广至仅具有线性尾部增长($\alpha \geq 1$)的非凸势函数?
  • RQ3光滑性阶数$\beta$与尾部增长$\alpha$之间如何相互作用,以决定LMC收敛速率的$\epsilon$-依赖性?
  • RQ4能否推导出一种依赖于矩的修正对数 Sobolev 不等式,以捕捉非对数凹目标下LMC的行为?
  • RQ5收敛速率是否以反映尾部增长结构的方式依赖于维度$d$,特别是在$\alpha=2$时?

主要发现

  • 对于具有$\alpha$-增长尾部($\alpha \in [1,2]$)且梯度为$\beta$-Hölder连续的势函数,LMC在KL散度上的收敛速率为$\widetilde{\mathcal{O}}\big(d^{\frac{1}{\beta} + \frac{1+\beta}{\beta}(\frac{2}{\alpha} - \mathbbm{1}_{\{\alpha \neq 1\}})} \epsilon^{-\frac{1}{\beta}}\big)$。
  • 只要$\alpha \geq 1$,速率的$\epsilon$-依赖性仅取决于光滑性阶数$\beta$,而与尾部增长速率$\alpha$无关,这将$\beta=1$时的最佳已知速率推广至非凸势函数。
  • 当$\alpha=2$时,维度依赖性恢复了$\beta=1$时的最佳已知$\widetilde{\mathcal{O}}(d\epsilon^{-1})$速率,即使对无穷远处退化凸势函数也成立。
  • 已建立具有显式常数的依赖于矩的修正对数 Sobolev 不等式,使得在弱假设下可推导出收敛速率。
  • 该框架适用于任意$\beta \in (0,1]$,包括非Lipschitz梯度,并允许凸势函数的有限扰动,因此覆盖了广泛的非凸目标。
  • 结果对LMC的最后一个迭代是紧的,提示高阶矩控制在高维、非对数凹采样中实现最优收敛至关重要。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。