Skip to main content
QUICK REVIEW

[論文レビュー] Primal Dual Interpretation of the Proximal Stochastic Gradient Langevin Algorithm

Adil Salim, Peter Richtárik|arXiv (Cornell University)|Jun 16, 2020
Markov Chains and Monte Carlo Methods参考文献 27被引用数 6
ひとこと要約

本稿は、滑らかで凸な関数と滑らかでない凸関数(可能性として無限大をとる)の和である合成ポテンシャルを持つ対数凸分布からのサンプリングに対して、Proximal Stochastic Gradient Langevin Algorithm (PSGLA) の原始双対解釈を提供する。Wasserstein空間における強い双対性を活用することで、強い凸性のもとで2-Wasserstein距離における${\mathcal{O}}(1/\varepsilon^{2})$の複雑度境界を確立し、これはFが凸かつ滑らかである場合のProjected Langevin Algorithmの${\mathcal{O}}(1/\varepsilon^{12})$の境界と比べ顕著に改善されている。

ABSTRACT

We consider the task of sampling with respect to a log concave probability distribution. The potential of the target distribution is assumed to be composite, extit{i.e.}, written as the sum of a smooth convex term, and a nonsmooth convex term possibly taking infinite values. The target distribution can be seen as a minimizer of the Kullback-Leibler divergence defined on the Wasserstein space ( extit{i.e.}, the space of probability measures). In the first part of this paper, we establish a strong duality result for this minimization problem. In the second part of this paper, we use the duality gap arising from the first part to study the complexity of the Proximal Stochastic Gradient Langevin Algorithm (PSGLA), which can be seen as a generalization of the Projected Langevin Algorithm. Our approach relies on viewing PSGLA as a primal dual algorithm and covers many cases where the target distribution is not fully supported. In particular, we show that if the potential is strongly convex, the complexity of PSGLA is $O(1/\varepsilon^2)$ in terms of the 2-Wasserstein distance. In contrast, the complexity of the Projected Langevin Algorithm is $O(1/\varepsilon^{12})$ in terms of total variation when the potential is convex.

研究の動機と目的

  • 対数凸分布からのサンプリングにおけるProximal Stochastic Gradient Langevin Algorithm (PSGLA) を分析するためのプライマルデュアルフレームワークを提供すること。
  • Wasserstein空間におけるKullback-Leibler発散の最小化に強い双対性を確立し、PSGLAの複雑度解析を可能にすること。
  • 2-Wasserstein距離の観点から、特にターゲット分布が完全にサポートされていない場合にPSGLAの非漸近的収束速度を導出すること。
  • ポテンシャル関数$G$が無限大の値をとる場合の複雑度境界を拡張し、制約付きサンプリング問題をカバーすること。

提案手法

  • 滑らかな凸関数$F$と滑らかでない凸関数$G$からなる合成ポテンシャル$V = F + G$を用いて、Wasserstein空間におけるKL発散の最小化としてサンプリング問題を定式化すること。
  • KL最小化問題に対する強い双対性を確立し、双対変数と$G$の共役関数を含むラグランジュ関数を導入すること。
  • PSGLAをプライマルデュアルアルゴリズムとして分析し、プライマル反復$x^{k+1}$をproximalステップで更新し、デュアル反復$y^{k+1}$をデュアル問題から導出すること。
  • 現在の分布とターゲット分布間の2-Wasserstein距離を含む再帰不等式を導出する。この不等式には双対ギャップ項とノイズ分散が含まれる。
  • 強い凸性を有する$F$と$G^*$の滑らかさを活用し、双対ギャップを用いてWasserstein距離の減衰を制御すること。
  • 確率的勾配とノイズの集中およびモーメントバウンドを適用し、再帰式における残差項を制御すること。

実験結果

リサーチクエスチョン

  • RQ1PSGLAは、合成対数凸分布からのサンプリングの文脈において、プライマルデュアルアルゴリズムとして解釈可能か?
  • RQ2Gが滑らかでない、あるいは無限大の値をとる場合に、PSGLAの2-Wasserstein距離における非漸近的収束速度は何か?
  • RQ3Total variationとWasserstein距離の観点から、PSGLAの複雑度はProjected Langevin Algorithmと比べてどのように異なるか?
  • RQ4Wasserstein空間における強い双対性を活用することで、制限付きサポートを持つサンプリングアルゴリズムのよりタイトな収束境界を導出可能か?

主な発見

  • ポテンシャル$F$が強く凸で、$G$が$1/\lambda_{G^*}$-滑らかである場合、PSGLAは2-Wasserstein距離において${\mathcal{O}}(1/\varepsilon^{2})$の複雑度境界を達成する。
  • これは、Fが凸かつ滑らかである場合のProjected Langevin AlgorithmのTotal variationにおける${\mathcal{O}}(1/\varepsilon^{12})$の境界と比べ顕著に優れている。
  • 強い双対性の結果から生じる双対ギャップが、Wasserstein距離の減衰を制御する再帰不等式を導出するのにも用いられている。
  • ある$x$に対して$G(x) = +\infty$であっても、解析は成り立つため、凸体上にサポートを持つターゲット分布のケースをカバーしている。
  • 適切な仮定の下で、追加の正則化子を含むStochastic Proximal Langevin Algorithm (SPLA) に対しても同様の収束速度が維持される。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。