Skip to main content
QUICK REVIEW

[論文レビュー] Minimizing Regret in Bandit Online Optimization in Unconstrained and Constrained Action Spaces.

Tatiana Tatarenko, Maryam Kamgarpour|arXiv (Cornell University)|Jun 13, 2018
Advanced Bandit Algorithms Research参考文献 23被引用数 3
ひとこと要約

本稿では、特徴的な確率的乱周波数スケーリング手法を用いた1点勾配推定を活用することで、制約なしおよび制約ありの行動空間の両方で、O(nT^{2/3})のレグレットレートを達成する、新しいゼロオーダーのオンライン凸最適化アルゴリズムを提案する。この手法は2点フィードバックに拡張され、理論的レグレット下界と一致する。

ABSTRACT

We consider online convex optimization with a zero-order oracle feedback. In particular, the decision maker does not know the explicit representation of the time-varying cost functions, or their gradients. At each time step, she observes the value of the cost function evaluated at her chosen action. The objective is to minimize the regret, that is, the difference between the sum of the costs she accumulates and that of the static optimal action had she known the sequence of cost functions a priori. We present a novel algorithm to minimize the regret in both unconstrained and constrained action spaces. Our algorithm hinges on a classical idea of one-point estimation of the gradients of the cost functions based on their observed values. However, our choice of the randomization introduced and consequently the proof techniques differ from those of past work. Letting T denote the number of queries of the zero-order oracle and n the problem dimension, the regret rate achieved is O(nT^{2/3}) for both constrained and unconstrained action spaces. Moreover, we adapt the presented algorithm to the setting with two-point feedback and demonstrate that the adapted procedure achieves the theoretical lower bound on the regret.

研究の動機と目的

  • 勾配情報が明示的に入手できないゼロオーダーのオракルフィードバックを伴うオンライン凸最適化を扱う。
  • このフィードバックモデル下で、制約なしおよび制約ありの行動空間において、レグレットを最小化する。
  • 従来の手法よりも優れた新しい乱周波数戦略を用いた勾配推定技術を開発する。
  • 提案手法が2点フィードバックに適応された場合、理論的レグレット下界に達することを示す。

提案手法

  • 関数値観測のみに基づいて、時間的に変化する目的関数の勾配を近似するため、確率的摂動を用いた1点勾配推定を採用する。
  • 従来の研究とは異なる、特定の乱周波数分布を導入し、よりタイトなレグレット解析を可能にする。
  • 制約付きの行動空間においても実行可能性を維持しながら、有界なレグレットを保証する投影を不要とする更新則を設計する。
  • 対称的摂動を用いることで、2点フィードバックに適応し、勾配推定の精度を向上させる。
  • 選択した乱周波数に特化した新しい集中不等式およびマルティングルールの議論を用いて、O(nT^{2/3})のレグレットバウンドを導出する。
  • 2点フィードバックバージョンが、既知の理論的レグレット下界に達することを確立し、その設定における最適性を確認する。

実験結果

リサーチクエスチョン

  • RQ1勾配情報が入手不可である状況下で、ゼロオーダーのオンライン最適化アルゴリズムが、制約なしおよび制約ありの両方の行動空間において、非線形なレグレットレートを達成できるか?
  • RQ21点勾配推定における乱周波数の選択が、高次元設定におけるレグレットバウンドにどのように影響するか?
  • RQ3提案手法を2点フィードバックに適応することで、レグレットの理論的下界に一致させられるか?
  • RQ4ゼロオーダーのオンライン凸最適化において、推定精度とレグレット成長の最適なトレードオフは何か?

主な発見

  • 提案手法は、ゼロオーダーのフィードバック下で、制約なしおよび制約ありの両方の行動空間において、O(nT^{2/3})のレグレットレートを達成する。
  • レグレットバウンドは、制約集合の具体的な構造に依存せず、次元数と時間枠にのみ依存する。
  • アルゴリズムの乱周波数戦略により、集中性の向上が実現され、従来手法に比べてよりタイトなレグレット解析が可能になる。
  • 2点フィードバックに適応した場合、アルゴリズムのレグレットは既知の理論的下界と一致し、その設定における最適性が確認される。
  • アルゴリズムは投影操作や目的関数の勾配の明示的知識を必要とせず、関数値クエリのみに依存する。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。