Skip to main content
QUICK REVIEW

[論文レビュー] A Modern Introduction to Online Learning

Francesco Orabona|arXiv (Cornell University)|Dec 31, 2019
Advanced Bandit Algorithms Research参考文献 92被引用数 68
ひとこと要約

オンライン凸最適化についての現代的で総合的なテキストで、オンライン学習アルゴリズム(FTL、OGD、OMD、FTRL)を後悔分析、適応的手法、そしてより広いトピックへの基礎的な結びつきとともに詳述する。

ABSTRACT

In this monograph, I introduce the basic concepts of Online Learning through a modern view of Online Convex Optimization. Here, online learning refers to the framework of regret minimization under worst-case assumptions. I present first-order and second-order algorithms for online learning with convex losses, in Euclidean and non-Euclidean settings. All the algorithms are clearly presented as instantiation of Online Mirror Descent or Follow-The-Regularized-Leader and their variants. Particular attention is given to the issue of tuning the parameters of the algorithms and learning in unbounded domains, through adaptive and parameter-free online learning algorithms. Non-convex losses are dealt through convex surrogate losses and through randomization. The bandit setting is also briefly discussed, touching on the problem of adversarial and stochastic multi-armed bandits. These notes do not require prior knowledge of convex analysis and all the required mathematical tools are rigorously explained. Moreover, all the included proofs have been carefully chosen to be as simple and as short as possible.

研究の動機と目的

  • オンライン学習フレームワークと後悔最小化を中核の目的として導入する。
  • オンライン勾配降下法、サブ勾配降下法、ミラー降下法、FTRL などの主要なオンラインアルゴリズムとそれらの後悔分析を提示する。
  • 適応的/パラメータフリー手法、強凸性、バンディット、非定常設定などの拡張を探る。
  • 凸分析、確率的最適化、学習理論など、オンライン学習をより広いトピックと関連づける。

提案手法

  • 任意の比較基準と敵対的な損失列に対して後悔を定義する。
  • コアアルゴリズム: Online Subgradient Descent、projection を伴う Online Gradient Descent、Follow-the-Regularized-Leader の派生を開発・分析する。
  • Online Mirror Descent を導入し、サブグラデントと Bregman ダイバージェンスとの関連を説明する。
  • 凸性と有界勾配の下でサブ線形の後悔を確立するための証明と補題(Be-the-Leader など)を提供する。

実験結果

リサーチクエスチョン

  • RQ1凹性を持つ(微分可能または非微分可能な)損失に対して、オンライン学習で保証される後悔界はどのようなものか?
  • RQ2さまざまな凸性と領域条件の下で、Online Gradient Descent、Subgradient Descent、Mirror Descent はどう比較されるか?
  • RQ3適応性およびパラメータフリー手法は、無界または複雑な領域でサブ線形後悔を達成する上でどのような役割を果たすか?
  • RQ4バンディット、鞍点問題、逐次投資などの高度な設定へオンライン学習の概念をどのように拡張できるか?

主な発見

  • FTL は敵対的な設定でサブ最適になり得るため、projection を用いた Online Gradient Descent が動機づけられる。
  • 射影付きオンライン勾配降下法は、凸微分可能な損失と有界領域の下でサブ線形後悔境界をもたらす。
  • Be-the-Leader 補題は、適応的リーダーの累積損失を固定の比較基準と関連づけることでサブ線形後悔を裏付ける。
  • Online Mirror Descent は Bregman ダイバージェンスを介して統一的な見方を提供し、非ユークリッド幾何も扱える。
  • 適応的・パラメータフリーの派生、および強凸性は性能を向上させ、適用範囲を広げる。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。