Skip to main content
QUICK REVIEW

[論文レビュー] On the Complexity Analysis of the Primal Solutions for the Accelerated Randomized Dual Coordinate Ascent

Huan Li, Zhouchen Lin|arXiv (Cornell University)|Jul 1, 2018
Stochastic Gradient Optimization Techniques参考文献 41被引用数 6
ひとこと要約

本稿は、加速型ランダム双対座標降下法(ARDCA)における原問題解の最適な $ O\left(\sqrt{\frac{n}{\epsilon}}\right) $ 迴数複雑性を確立し、双対収束レートと一致させた。二次的機能的成長条件の下で線形収束を証明し、従来のスムージング法やCatalystに基づく手法で見られる $ \log(1/\epsilon) $ 要因を排除し、正則化された経験的リスク最小化における双対座標降下法で初めて原問題の最適な複雑性を達成した。

ABSTRACT

Dual first-order methods are essential techniques for large-scale constrained convex optimization. However, when recovering the primal solutions, we need $T(ε^{-2})$ iterations to achieve an $ε$-optimal primal solution when we apply an algorithm to the non-strongly convex dual problem with $T(ε^{-1})$ iterations to achieve an $ε$-optimal dual solution, where $T(x)$ can be $x$ or $\sqrt{x}$. In this paper, we prove that the iteration complexity of the primal solutions and dual solutions have the same $O\left(\frac{1}{\sqrtε} ight)$ order of magnitude for the accelerated randomized dual coordinate ascent. When the dual function further satisfies the quadratic functional growth condition, by restarting the algorithm at any period, we establish the linear iteration complexity for both the primal solutions and dual solutions even if the condition number is unknown. When applied to the regularized empirical risk minimization problem, we prove the iteration complexity of $O\left(n\log n+\sqrt{\frac{n}ε} ight)$ in both primal space and dual space, where $n$ is the number of samples. Our result takes out the $\left(\log \frac{1}ε ight)$ factor compared with the methods based on smoothing/regularization or Catalyst reduction. As far as we know, this is the first time that the optimal $O\left(\sqrt{\frac{n}ε} ight)$ iteration complexity in the primal space is established for the dual coordinate ascent based stochastic algorithms. We also establish the accelerated linear complexity for some problems with nonsmooth loss, i.e., the least absolute deviation and SVM.

研究の動機と目的

  • 加速型ランダム双対座標降下法(ARDCA)における双対と原問題の収束複雑性のギャップを埋めること。
  • 非強い凸問題における双対反復から得られる原問題解の反復複雑性を分析すること。
  • 正則化された経験的リスク最小化(ERM)における原問題解の最適な $ O\left(\sqrt{\frac{n}{\epsilon}}\right) $ 反復複雑性を確立すること。
  • 二次的機能的成長条件の下で、原問題と双対問題の両方の解が線形収束することを示すこと、条件数を事前に知らなくてもよいこと。
  • スムージング法やCatalystに基づくアプローチで一般的に見られる $ \log(1/\epsilon) $ 要因を、双対座標降下法の原問題複雑性から排除すること。

提案手法

  • ラグランジュ双対枠組みと双対ギャップを用いて、双対反復から得られる原問題解の回復を分析する。
  • 原問題誤差が $ O\left(\sqrt{D(\mathbf{u}^K) - D(\mathbf{u}^*)} + D(\mathbf{u}^K) - D(\mathbf{u}^*)\right) $ で有界であることを証明し、原問題と双対問題の収束を結びつける。
  • 双対関数が二次的機能的成長条件を満たす場合に線形収束を達成するため、定期的な再起動戦略を導入する。
  • リャプノフ関数と再帰的解析を用いて、期待される双対ギャップと原問題の部分最適性をバインドする。
  • $ n $ 個のサンプルを持つERM問題にこの手法を適用し、原問題および双対空間において $ O\left(n\log n + \sqrt{\frac{n}{\epsilon}}\right) $ の複雑性を導出する。
  • 最小絶対偏差やSVMのような非滑らかな損失関数に対しても、加速型線形収束を確立する。

実験結果

リサーチクエスチョン

  • RQ1ARDCAの原問題解の収束複雑性は、双対収束複雑性と一致させられるか?
  • RQ2二次的機能的成長条件が、ARDCAにおける原問題と双対問題の両方の線形収束を可能にするか?
  • RQ3スムージング法やCatalystに基づくアプローチで一般的に見られる $ \log(1/\epsilon) $ 要因は、双対座標降下法の原問題複雑性から排除可能か?
  • RQ4双対座標降下法を用いた正則化された経験的リスク最小化問題における、$ O\left(\sqrt{\frac{n}{\epsilon}}\right) $ の複雑性は最適か?
  • RQ5ARDCAは、L1やSVMのような非滑らかな損失関数に対しても、加速型線形収束を達成できるか?

主な発見

  • ARDCAの原問題解の収束複雑性は $ O\left(\frac{1}{\sqrt{\epsilon}}\right) $ であり、双対複雑性と一致し、長年のギャップを解消した。
  • 二次的機能的成長条件の下では、条件数を事前に知らなくても、原問題と双対問題の両方が線形収束を達成する。
  • 正則化された経験的リスク最小化において、反復複雑性は原問題および双対空間で $ O\left(n\log n + \sqrt{\frac{n}{\epsilon}}\right) $ である。
  • スムージング法やCatalystに基づく手法と比較して、$ \log(1/\epsilon) $ 要因が排除され、双対座標降下法における初めての最適な $ O\left(\sqrt{\frac{n}{\epsilon}}\right) $ 原問題複雑性を達成した。
  • 機能的成長条件の下で、最小絶対偏差やSVMを含む非滑らかな損失関数を有する問題に対しても線形収束が確立された。
  • 双対ギャップが原問題の部分最適性を制御することを解析で証明し、再帰的リャプノフ関数技術によりタイトなバインドが可能であることを示した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。