[論文レビュー] Stability Conditions for Online Learnability
この論文は、一般の学習設定における敵対的データ系列下で、オンライン安定性——一様LOO安定性の変種——が、レギュレータリゼーション付き経験的リスク最小化(RERM)型アルゴリズムにおいて、no-regretオンライン学習を達成するのに十分であることを確立している。FTRL、ミラー降下、およびランダム化手法(例:Hedge)が、それらのバックエンドバッチ学習アルゴリズムが一様オンライン安定性を満たす場合にno-regretを達成することを証明し、二値分類においてこの条件が十分かつ必要であることを示している。
Stability is a general notion that quantifies the sensitivity of a learning algorithm's output to small change in the training dataset (e.g. deletion or replacement of a single training sample). Such conditions have recently been shown to be more powerful to characterize learnability in the general learning setting under i.i.d. samples where uniform convergence is not necessary for learnability, but where stability is both sufficient and necessary for learnability. We here show that similar stability conditions are also sufficient for online learnability, i.e. whether there exists a learning algorithm such that under any sequence of examples (potentially chosen adversarially) produces a sequence of hypotheses that has no regret in the limit with respect to the best hypothesis in hindsight. We introduce online stability, a stability condition related to uniform-leave-one-out stability in the batch setting, that is sufficient for online learnability. In particular we show that popular classes of online learners, namely algorithms that fall in the category of Follow-the-(Regularized)-Leader, Mirror Descent, gradient-based methods and randomized algorithms like Weighted Majority and Hedge, are guaranteed to have no regret if they have such online stability property. We provide examples that suggest the existence of an algorithm with such stability condition might in fact be necessary for online learnability. For the more restricted binary classification setting, we establish that such stability condition is in fact both sufficient and necessary. We also show that for a large class of online learnable problems in the general learning setting, namely those with a notion of sub-exponential covering, no-regret online algorithms that have such stability condition exists.
研究の動機と目的
- 敵対的データ系列下での一般学習設定におけるオンライン学習可能性の十分条件を同定すること。
- i.i.d.バッチ設定における安定性に基づく学習可能性理論を、オンラインで敵対的な設定へと拡張すること。
- オンライン安定性が、制限付き設定におけるオンライン学習可能性にとって十分であるだけでなく、必要であるかどうかも同定すること。
- 仮説空間のカバー性、特に指数的カバー性(sub-exponential covering)とオンライン学習可能性を結びつけること。
- 二値分類において、オンライン学習可能性と一様オンライン安定性または一様LOO安定性を持つRERMアルゴリズムの存在が同値であることを確立すること。
提案手法
- オンライン学習に適応した一様LOO安定性に類似した安定性条件として、オンライン安定性を導入する。
- FTRL、ミラー降下、勾配ベースの手法を、正則化経験的リスク最小化(RERM)の具体例として分析する。
- 仮説空間の有限$ε$-カバーにHedgeアルゴリズムを適用し、レジットバウンドを導出する。
- 一般問題におけるオンライン学習可能性を特徴付けるために、指数的カバー性の概念を用いる。
- 次の形のレジットバウンドを導出する:$\epsilon_{\text{regret}}(t) \leq B\sqrt{2\log(N(\mathcal{H},\mathcal{Z},f,\epsilon_m))}\left[\frac{3}{\sqrt{t}} + \frac{\log t}{2t} + \frac{1+2\ln 2}{2t}\right] + \epsilon_m$($t \leq m$ に対して)。
- Lipschitz性および有界直径の仮定の下で、$\u03b5$-カバー上のランダム化RERMアルゴリズムが、$O(\sqrt{\log m / t})$のレジットを達成することを示す。
実験結果
リサーチクエスチョン
- RQ1オンライン安定性は、一般学習設定下でno-regretオンライン学習を達成するのに十分か?
- RQ2二値分類において、オンライン学習可能性に対してオンライン安定性が必要であることが示せるか?
- RQ3指数的カバー性を持つすべてのオンライン学習可能な問題は、一様オンライン安定RERMを介してno-regret学習が可能か?
- RQ4一般設定下で、一様オンライン安定または一様LOO安定RERMアルゴリズムの存在が、オンライン学習可能性を特徴づけるか?
- RQ5一般学習フレームワーク下で、指数的カバー性の概念とオンライン学習可能性が同値であるか?
主な発見
- RERM型アルゴリズムにおいて、オンライン安定性は一般オンライン学習設定下でno-regret学習を達成するのに十分である。
- 二値分類において、オンライン学習可能性は、(おそらくランダム化された)一様オンライン安定RERMアルゴリズムの存在と同値である。
- 指数的カバー性を持つ問題では、ランダム化された一様LOO安定RERMアルゴリズムを介してno-regretオンライン学習が達成可能である。
- このようなアルゴリズムのレジットバウンドは、$\epsilon_{\text{regret}}(t) \leq B\sqrt{2\log(N(\mathcal{H},\mathcal{Z},f,\epsilon_m))}\left[\frac{3}{\sqrt{t}} + \frac{\log t}{2t} + \frac{1+2\ln 2}{2t}\right] + \epsilon_m$($t \leq m$ に対して)である。
- Lipschitz性および有界直径の仮定の下で、レジットレートは$O\left(\sqrt{\frac{\log(K) + d\log(mD)}{t}}\right)$であり、$t \to \infty$のときno-regretを達成する。
- 連続でない損失関数(例:有理数/無理数の指標損失)に対しても、有限$\u03b5$-カバー上でHedgeを適用することで、$O(1/\sqrt{m})$のレジットが達成可能である。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。