Skip to main content
QUICK REVIEW

[論文レビュー] Input Perturbations for Adaptive Regulation and Learning.

Mohamad Kazem Shirani Faradonbeh, Ambuj Tewari|arXiv (Cornell University)|Nov 10, 2018
Advanced Bandit Algorithms Research被引用数 12
ひとこと要約

本論文は、MIMO線形システムにおける適応的制御および推定のため、時間に関して(ほぼ)平方根オーダーの非漸近的リグレットバウンドを達成するために、入力信号の摂動を用いた摂動付きグリーディー方策を提案する。マルティングゲール理論と方策分解を活用することで、システムパラメータの事前知識を必要とせず、計算的に非実行可能または不安定な既存の手法の限界を克服し、リグレットと学習精度の高確率バウンドを達成する。

ABSTRACT

Design of adaptive algorithms for simultaneous regulation and estimation of MIMO linear dynamical systems is a canonical reinforcement learning problem. Efficient policies whose regret (i.e. increase in the cost due to uncertainty) scales at a square-root rate of time have been studied extensively in the recent literature. Nevertheless, existing strategies are computationally intractable and require a priori knowledge of key system parameters. The only exception is a randomized Greedy regulator, for which asymptotic regret bounds have been recently established. However, randomized Greedy leads to probable fluctuations in the trajectory of the system, which renders its finite time regret suboptimal. This work addresses the above issues by designing policies that utilize input signals perturbations. We show that perturbed Greedy guarantees non-asymptotic regret bounds of (nearly) square-root magnitude w.r.t. time. More generally, we establish high probability bounds on both the regret and the learning accuracy under arbitrary input perturbations. The settings where Greedy attains the information theoretic lower bound of logarithmic regret are also discussed. To obtain the results, state-of-the-art tools from martingale theory together with the recently introduced method of policy decomposition are leveraged. Beside adaptive regulators, analysis of input perturbations captures key applications including remote sensing and distributed control.

研究の動機と目的

  • MIMO線形システムの既存の適応的制御アルゴリズムにおける計算の非実行可能性と、システムパラメータの事前知識の欠如という問題に対処すること。
  • 確率的グリーディー方策の有限時間リグレットの非最適性および軌道のフラクチュエーションを克服すること。
  • 任意の入力摂動に対して、リグレットおよび学習精度の高確率バウンドを確立すること。
  • グリーディー方策が情報理論的対数的リグレット下界に達する条件を分析すること。
  • 入力摂動解析を通じて、リモートセンシングおよび分散制御への適用可能性を拡張すること。

提案手法

  • 制御された入力摂動を注入することで、探索性と安定性を向上させる摂動付きグリーディー方策の設計。
  • 最新のマルティングゲール理論を適用し、リグレットおよび推定誤差の高確率集中バウンドを導出する。
  • 最近導入された方策分解法を活用し、探索と活用のコンponentsを分離する。
  • リグレットおよび学習精度のバウンドを、時間Tに関して(ほぼ)√Tのスケーリングで定式化する。
  • 任意の入力摂動がシステム軌道および収束特性に与える影響を分析する。
  • グリーディー方策が情報理論的対数的リグレット下界に達する条件を確立する。

実験結果

リサーチクエスチョン

  • RQ1入力摂動を用いることで、MIMO適応的制御における非漸近的リグレットバウンドを√Tオーダーで達成できるか?
  • RQ2任意の入力摂動が、リグレットおよび学習精度の高確率バウンドにどのように影響するか?
  • RQ3グリーディー方策が情報理論的対数的リグレット下界に達する条件は何か?
  • RQ4提案手法は、システムパrameterの事前知識を必要とせず、計算的に実行可能であるか?
  • RQ5方策分解およびマルティングゲールツールは、有限時間性能保証を導出する上で果たす役割は何か?

主な発見

  • 摂動付きグリーディー方策は、確率的グリーディーの有限時間非最適性を上回り、時間に関して(ほぼ)平方根オーダーの非漸近的リグレットバウンドを達成する。
  • 任意の入力摂動に対して、リグレットおよび学習精度の両方の高確率バウンドが確立され、頑健な性能を保証する。
  • 本手法は計算的に実行可能であり、大多数の既存手法とは異なり、主要なシステムパラメータの事前知識を必要としない。
  • 理論的分析により、特定の条件下でグリーディー方策が情報理論的対数的リグレット下界に達することが確認された。
  • 本フレームワークは適応制御を越えて、リモートセンシングおよび分散制御システムへの応用が可能である。
  • 方策分解とマルティングゲール理論の統合により、適応的制御におけるタイトな有限時間性能保証が可能となった。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。