Skip to main content
QUICK REVIEW

[論文レビュー] Optimal Decentralized Distributed Algorithms for Stochastic Convex Optimization

Eduard Gorbunov, Darina Dvinskikh|arXiv (Cornell University)|Nov 17, 2019
Stochastic Gradient Optimization Techniques参考文献 93被引用数 34
ひとこと要約

本稿は、非正確な近似射影ステップと勾配スライディングを用いた、プライマルおよびデュアルアプローチを活用して、アフィン制約付きの確率的凸最適化のための新規分散型分散アルゴリズムを提案する。通信ラウンド数とオракル呼び出し回数の観点から最適収束レートを確立し、ε-精度を達成するためのO(√L/μ log(1/ε))回の反復とeO(max{√L/μ, σ²/ε})回のオラクル呼び出しを達成する。バイアスありおよびバイアスなしの確率的オラクルにおいて、タイトな境界を示す。

ABSTRACT

We consider stochastic convex optimization problems with affine constraints and develop several methods using either primal or dual approach to solve it. In the primal case we use special penalization technique to make the initial problem more convenient for using optimization methods. We propose algorithms to solve it based on Similar Triangles Method with Inexact Proximal Step for the convex smooth and strongly convex smooth objective functions and methods based on Gradient Sliding algorithm to solve the same problems in the non-smooth case. We prove the convergence guarantees in smooth convex case with deterministic first-order oracle. We propose and analyze three novel methods to handle stochastic convex optimization problems with affine constraints: SPDSTM, R-RRMA-AC-SA and SSTM_sc. All methods use stochastic dual oracle. SPDSTM is the stochastic primal-dual modification of STM and it is applied for the dual problem when the primal functional is strongly convex and Lipschitz continuous on some ball. R-RRMA-AC-SA is an accelerated stochastic method based on the restarts of RRMA-AC-SA and SSTM_sc is just stochastic STM for strongly convex problems. Both methods are applied to the dual problem when the primal functional is strongly convex, smooth and Lipschitz continuous on some ball and use stochastic dual first-order oracle. We develop convergence analysis for these methods for the unbiased and biased oracles respectively. Finally, we apply all aforementioned results and approaches to solve decentralized distributed optimization problem and discuss optimality of the obtained results in terms of communication rounds and number of oracle calls per node.

研究の動機と目的

  • アフィン制約下での最適分散型分散アルゴリズムを、確率的凸最適化のために開発すること。
  • 分散環境における理論的収束レートと実用的通信効率のギャップを埋めること。
  • 特に分散ネットワークにおけるバイアスありおよびバイアスなしの確率的オラクルの下での収束を分析すること。
  • 滑らかで強く凸な問題における通信ラウンド数とオラクル呼び出し回数のタイトな境界を確立すること。
  • 非正確なオラクルとバイアスのある確率的勾配を伴う加速手法をプライマル・デュアルフレームワークに拡張すること。

提案手法

  • 類似三角形法(STM)と非正確な近似射影ステップを用いて、効率的な最適化が可能なように、ペナルティ技法を用いて問題を再定式化するプライマルアプローチを提案する。
  • 非滑らかケースのための勾配スライディングアルゴリズムを導入し、確率的サブ勾配の効率的取り扱いを可能にする。
  • バイアスありおよびバイアスなしの確率的オラクルに特化した、3つの新規手法(SPDSTM, SSTM sc, R-RRMA-AC-SA2)を開発する。
  • 強い凸性を持つデュアル関数とリスタート技法を用いて、収束を加速するデュアルアプローチを適用する。
  • デュアル問題に対して直接加速を適用し、リャプノフ関数と確率的誤差バウンドを用いて収束を証明する。
  • 分散還元と非正確なオラクル解析を用いて、反復回数およびオラクル複雑度のタイトな境界を導出する。

実験結果

リサーチクエスチョン

  • RQ1アフィン制約付きの分散型確率的凸最適化における最適通信複雑度は何か?
  • RQ2分散ネットワークにおけるバイアスのある確率的オラクルに対して、加速されたプライマル・デュアル手法をどのように設計できるか?
  • RQ3分散型確率的最適化において、オラクル呼び出し回数と通信ラウンド数のトレードオフは何か?
  • RQ4非正確な近似射影ステップと勾配スライディングを組み合わせることで、非滑らか設定において最適収束を達成できるか?
  • RQ5バイアスありおよびバイアスなしの確率的オラクルは、分散型デュアル手法における収束レートにどのように影響するか?

主な発見

  • 提案されたR-RRMA-AC-SA2は、バイアスなしの場合にε-精度を達成するためのO(√L/μ log(1/ε))通信ラウンドで最適収束を達成する。
  • バイアスありオラクル(SPDSTM, SSTM sc)の場合、O(√L/μ log(1/ε))の収束を達成し、eO(max{√L/μ, σ²/ε})回のオラクル呼び出しを要する。
  • アルゴリズムフレームワークにより、高確率でf(˜xN) − f(x*) ≤ (2 + √(2C/λmax(A⊤A) + G1))ε/Ryが成り立つことが保証される。
  • 確率が1 − 4β以上である確率で、∥A˜xN∥₂ ≤ (1 + √(2C) + G1/√λmax(A⊤A))ε/Ryが成立する。
  • 解析により、ε-精度を達成するためのオラクル呼び出し総数がeO(max{√L/μ, σ²/ε})であることが示され、既知の下界と一致する。
  • 通信ラウンド数およびオラクル複雑度の両面で収束レートが最適であり、確率的リャプノフ関数と非正確なオラクル解析を用いてタイトな境界が導出された。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。