Skip to main content
QUICK REVIEW

[論文レビュー] Chi-square Tests Driven Method for Learning the Structure of Factored MDPs

Thomas Degris, Olivier Sigaud|arXiv (Cornell University)|Jun 27, 2012
Reinforcement Learning in Robotics参考文献 13被引用数 11
ひとこと要約

本稿では、確率分布における統計的独立性を検出するためにカイ二乗検定を用いることにより、要因分解されたマーカフ・意思決定過程(FMDP)の構造を段階的に学習する手法SPITIを提案する。カイ二乗検定のしきい値を調整することで、価値関数における相対誤差が小さいポリシーを生成するコンactなモデルを構築できる。FMDPフレームワーク内での効果的な一般化により、大規模な確率的強化学習問題において、表形式の手法を上回る性能を発揮する。

ABSTRACT

SDYNA is a general framework designed to address large stochastic reinforcement learning problems. Unlike previous model based methods in FMDPs, it incrementally learns the structure and the parameters of a RL problem using supervised learning techniques. Then, it integrates decision-theoric planning algorithms based on FMDPs to compute its policy. SPITI is an instanciation of SDYNA that exploits ITI, an incremental decision tree algorithm, to learn the reward function and the Dynamic Bayesian Networks with local structures representing the transition function of the problem. These representations are used by an incremental version of the Structured Value Iteration algorithm. In order to learn the structure, SPITI uses Chi-Square tests to detect the independence between two probability distributions. Thus, we study the relation between the threshold used in the Chi-Square test, the size of the model built and the relative error of the value function of the induced policy with respect to the optimal value. We show that, on stochastic problems, one can tune the threshold so as to generate both a compact model and an efficient policy. Then, we show that SPITI, while keeping its model compact, uses the generalization property of its learning method to perform better than a stochastic classical tabular algorithm in large RL problem with an unknown structure. We also introduce a new measure based on Chi-Square to qualify the accuracy of the model learned by SPITI. We qualitatively show that the generalization property in SPITI within the FMDP framework may prevent an exponential growth of the time required to learn the structure of large stochastic RL problems.

研究の動機と目的

  • 大規模な要因分解MDP(FMDP)の構造を、確率的強化学習環境で学習する課題に対処すること。
  • 教師あり学習手法を用いて、構造とパラメータを段階的に学習する手法を開発すること。
  • 得られたモデルがコンパクトなまま、高いポリシー性能を維持すること。
  • 段階的学習の一般化能力を活用し、学習時間の指数的増加を回避すること。
  • FMDP構造学習におけるモデルの正確さを評価するカイ二乗に基づく指標を導入すること。

提案手法

  • SPITIは、遷移関数および報酬関数における確率変数間の統計的独立性を検出するためにカイ二乗検定を用いる。
  • 段階的に、局所構造を持つ動的ベイジアンネットワークを構築して、FMDPの遷移ダイナミクスを表現する。
  • 学習済みモデルに基づいてポリシーを計算するために、構造的価値反復の段階的版を適用する。
  • カイ二乗検定のしきい値がモデルの複雑さを制御する:低いしきい値は密度の高いモデルを、高いしきい値はスパarsなモデルを生成する。
  • 学習済みモデル構造の正確さを評価するための新しいカイ二乗に基づく指標を導入する。
  • 報酬関数および条件付き確率分布を学習するために、ITI(段階的決定木アルゴリズム)を統合する。

実験結果

リサーチクエスチョン

  • RQ1大規模な確率的強化学習問題において、要因分解MDPの構造をどのように効率的に学習できるか?
  • RQ2カイ二乗検定のしきい値、モデルサイズ、価値関数誤差におけるポリシー品質との関係は何か?
  • RQ3一般化を活用することで、コンパクトなFMDPモデルを学習しつつ、高いポリシー性能を維持できるか?
  • RQ4カイ二乗駆動の構造学習手法は、大規模問題における古典的表形式Q学習と比べてどのように異なるか?
  • RQ5学習手法の一般化特性は、大規模FMDPにおける学習時間の指数的増加を防げるか?

主な発見

  • カイ二乗検定のしきい値を調整することで、モデルのコンパクトさとポリシーの正確さのトレードオフを実現でき、価値関数における相対誤差を低く抑えることができる。
  • SPITIは、構造が未知の大規模な確率的問題において、古典的表形式強化学習アルゴリズムを上回る性能を発揮する。
  • 一般化を活用することで学習効率を向上させつつ、モデルのコンパクトさを維持することができる。
  • カイ二乗に基づくモデル正確性指標は、ポリシー性能と相関するモデル品質の定性的な指標を提供する。
  • 段階的構造学習アプローチにより、学習時間の指数的増加を防止でき、大規模FMDPへのスケーラビリティを実現する。
  • 実験結果から、真の構造が未知であっても、FMDPフレームワーク内での効果的な一般化のおかげで、SPITIは効率的な学習が可能であることが示された。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。