[論文レビュー] Nonparametric General Reinforcement Learning
この論文は、確率的環境におけるThompsonサンプリングの漸近的最適性を確立し、マルチエージェント設定における「真実の粒」問題を解決することで、非パrametric一般強化学習を前進させる。事前分布に真実の粒が含まれる場合、Thompsonサンプリングは未知の計算可能マルチエージェント環境においてε-ナッシュ均衡に収束することを証明する一方で、AIXIの非計算可能性と、事前分布依存の最適性基準の限界も示している。
Reinforcement learning (RL) problems are often phrased in terms of Markov decision processes (MDPs). In this thesis we go beyond MDPs and consider RL in environments that are non-Markovian, non-ergodic and only partially observable. Our focus is not on practical algorithms, but rather on the fundamental underlying problems: How do we balance exploration and exploitation? How do we explore optimally? When is an agent optimal? We follow the nonparametric realizable paradigm. We establish negative results on Bayesian RL agents, in particular AIXI. We show that unlucky or adversarial choices of the prior cause the agent to misbehave drastically. Therefore Legg-Hutter intelligence and balanced Pareto optimality, which depend crucially on the choice of the prior, are entirely subjective. Moreover, in the class of all computable environments every policy is Pareto optimal. This undermines all existing optimality properties for AIXI. However, there are Bayesian approaches to general RL that satisfy objective optimality guarantees: We prove that Thompson sampling is asymptotically optimal in stochastic environments in the sense that its value converges to the value of the optimal policy. We connect asymptotic optimality to regret given a recoverability assumption on the environment that allows the agent to recover from mistakes. Hence Thompson sampling achieves sublinear regret in these environments. Our results culminate in a formal solution to the grain of truth problem: A Bayesian agent acting in a multi-agent environment learns to predict the other agents' policies if its prior assigns positive probability to them (the prior contains a grain of truth). We construct a large but limit computable class containing a grain of truth and show that agents based on Thompson sampling over this class converge to play Nash equilibria in arbitrary unknown computable multi-agent environments.
研究の動機と目的
- マルコフ決定過程を超える一般強化学習における根本的課題に取り組むこと、特に非マルコフ的で、非定常的かつ部分的に観測可能な環境において。
- ベイジアン強化学習エージェントの限界、特にAIXIの限界を調査し、Legg-Hutter知能のような事前分布依存の最適性基準の主観性を露呈すること。
- 一般強化学習設定におけるベイジアンエージェントに対して、客観的な最適性保証を確立すること、特にThompsonサンプリングを通じて。
- マルチエージェント環境における「真実の粒」問題を解決することを目的とし、真実の粒を含む極限計算可能クラスの事前分布を構築し、ε-ナッシュ均衡への収束を可能にする。
- アリトメティカル階層を用いて、普遍的強化学習エージェント(AIXIや知識探求型エージェントを含む)の計算可能性と非計算可能性を分析すること。
提案手法
- 非パrametric実現可能パラダイムを用いる:データが既知の可算クラスの候補内にある未知の源から来ると仮定する。
- Carathéodoryの拡張定理を適用して、有限文字列上の関数 q から無限列上の確率測度を構成する。
- 回復可能性の仮定の下で、値の収束とサブ線形レグレットの両方が成立することを示し、Thompsonサンプリングが確率的環境で漸近的に最適であることを確立する。
- AIXIが極限計算不能であり、アリトメティカル階層の高いレベルに位置することを示し、上界と下界による境界を用いて非計算可能性を証明する。
- 真実の粒を含む、広大で極限計算可能な環境クラスを構築し、マルチエージェント設定におけるε-ナッシュ均衡への収束を可能にする。
- σ-劣加法性と積位相におけるコンパクト性を含む測度論的道具を用いて、誘導される確率測度の存在と一意性を証明する。
実験結果
リサーチクエスチョン
- RQ1Thompsonサンプリングは、確率的で非マルコフ的かつ部分的に観測可能な環境において、漸近的最適性を達成できるか?
- RQ2AIXIの最適性はどの程度事前分布の選択に依存しており、その依存性は客観的に正当化可能か?
- RQ3任意の未知の計算可能なマルチエージェント環境において、ε-ナッシュ均衡に収束するベイジアン強化学習エージェントは存在するか?
- RQ4AIXIおよび関連する普遍的エージェントの非計算可能性の正確なレベルは、アリトメティカル階層の中でどこに位置するか?
- RQ5主観的な事前分布の選択に依存しない、一般強化学習において客観的な最適性保証を満たすベイジアンエージェントを構築可能か?
主な発見
- Thompsonサンプリングは、確率的環境において漸近的に最適であり、その価値は最適方策の価値に収束する。
- 回復可能性の仮定の下で、Thompsonサンプリングは確率的環境でサブ線形レグレットを達成する。
- AIXIは極限計算不能であり、アリトメティカル階層の高いレベルに位置し、根本的に非計算可能である。
- AIXIの極限計算可能なε-最適近似が存在し、実装可能な代替手段を提供する。
- すべての計算可能環境のクラスにおいて、いかなる方策もパレート最適であるため、AIXIの類似基準のような事前分布依存の最適性保証は根本的に崩れる。
- 真実の粒を含む極限計算可能な事前分布クラス上でThompsonサンプリングを用いるベイジアンエージェントは、任意の未知の計算可能なマルチエージェント環境において、ε-ナッシュ均衡への戦略的収束を達成する。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。