Skip to main content
QUICK REVIEW

[论文解读] High-dimensional Bayesian inference via the Unadjusted Langevin Algorithm

Alain Durmus, Éric Moulines|arXiv (Cornell University)|May 5, 2016
Markov Chains and Monte Carlo Methods被引用 10
一句话总结

本文在高维贝叶斯推断中为非调整的朗之万算法(ULA)提供了非渐近收敛界,建立了在强凸性和利普希茨梯度假设下,Wasserstein距离和总变差距离对维度的显式依赖关系。该研究改进了先前的界限,并针对固定精度或固定计算预算场景优化了步长与迭代次数。

ABSTRACT

We consider in this paper the problem of sampling a high-dimensional probability distribution $π$ having a density with respect to the Lebesgue measure on $\mathbb{R}^d$, known up to a normalization constant $x \mapsto π(x)= \mathrm{e}^{-U(x)}/\int_{\mathbb{R}^d} \mathrm{e}^{-U(y)} \mathrm{d} y$. Such problem naturally occurs for example in Bayesian inference and machine learning. Under the assumption that $U$ is continuously differentiable, $ abla U$ is globally Lipschitz and $U$ is strongly convex, we obtain non-asymptotic bounds for the convergence to stationarity in Wasserstein distance of order $2$ and total variation distance of the sampling method based on the Euler discretization of the Langevin stochastic differential equation, for both constant and decreasing step sizes. The dependence on the dimension of the state space of these bounds is explicit. The convergence of an appropriately weighted empirical measure is also investigated and bounds for the mean square error and exponential deviation inequality are reported for functions which are measurable and bounded. An illustration to Bayesian inference for binary regression is presented to support our claims.

研究动机与目标

  • 建立非调整朗之万算法(ULA)在高维设置下收敛到目标后验分布的非渐近收敛速率。
  • 推导出在Wasserstein距离和总变差距离上的显式边界,这些边界显式依赖于维度 $d$。
  • 分析恒定步长与递减步长两种情形,针对固定精度与固定计算预算进行优化。
  • 在强凸性和利普希茨梯度条件下,改进文献中关于总变差距离的现有界限。
  • 研究加权经验测度的收敛性,并提供均方误差与指数偏差的界。

提出的方法

  • ULA被表述为过阻尼朗之万SDE的Euler-Maruyama半离散化形式:$X_{k+1} = X_k - \gamma_{k+1} \nabla U(X_k) + \sqrt{2\gamma_{k+1}} Z_{k+1}$。
  • 在以下假设下分析收敛性:$U$ 连续可微,$\nabla U$ 全局利普希茨连续,且 $U$ 强凸。
  • 采用耦合技术来界定第 $n$ 次迭代与目标分布 $\pi$ 之间的总变差距离。
  • 通过反向归纳法分析耦合时间,利用高斯尾部概率界和 $V$-范数进行误差传播。
  • 利用 $\Xi_{k_1,k_2}$ 类型的方差和与步长序列,推导出Wasserstein距离和总变差距离的显式边界。
  • 分析涵盖固定步长与递减步长方案,对步长与迭代次数进行优化,以实现固定精度或固定计算预算。

实验结果

研究问题

  • RQ1在强凸性和利普希茨梯度条件下,ULA在Wasserstein距离和总变差距离中的非渐近收敛速率是什么?
  • RQ2对于固定步长与递减步长,收敛速率如何显式依赖于维度 $d$?
  • RQ3在相同假设下,能否推导出比先前工作更紧的总变差距离边界?
  • RQ4在总变差距离中实现固定精度 $\varepsilon$ 的情况下,步长与迭代次数之间最优权衡是什么?
  • RQ5ULA生成的加权经验测度在均方误差与指数偏差方面的收敛特性如何?

主要发现

  • 本文建立了ULA第 $n$ 次迭代与目标分布 $\pi$ 之间总变差距离的非渐近界,其在固定精度 $\varepsilon$ 下的尺度为 $\mathcal{O}(d\varepsilon^{-2})$,优于先前结果。
  • 对于固定步长 $\gamma$,在强凸性条件下,总变差距离的收敛速率被界为 $\mathcal{O}(d\gamma^{-1}e^{-\mu \gamma n})$,并显式体现了对 $d$ 的依赖。
  • 对于满足 $\sum_{k=1}^\infty \gamma_k = \infty$ 且 $\gamma_k \to 0$ 的递减步长,建立了在Wasserstein距离下趋于平稳分布的收敛性,且显式体现了维度依赖。
  • 加权经验测度的均方误差被界为 $\mathcal{O}(\gamma + \gamma^{1/2} \sqrt{d/n})$(对利普希茨函数),并提供了指数偏差不等式。
  • 分析表明,在总变差距离中实现 $\mathcal{O}(d\varepsilon^{-2})$ 的迭代复杂度时,无需使用预热启动,优于早期研究中所需的要求。
  • 在恒定步长下,$\pi$ 与 ULA 的平稳分布 $\pi_\gamma$ 之间距离的界在 $\gamma \to 0$ 时为 $\mathcal{O}(\gamma^{1/2})$,与已知的渐近行为一致。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。