Skip to main content
QUICK REVIEW

[论文解读] Nonparametric General Reinforcement Learning

Jan Leike|arXiv (Cornell University)|Nov 28, 2016
Multi-Criteria Decision Making被引用 6
一句话总结

该论文通过在随机环境中建立 Thompson 抽样渐近最优性,并解决多智能体环境中的‘真相颗粒’问题,推进了非参数通用强化学习。论文证明,当先验包含真相颗粒时,Thompson 抽样在未知可计算的多智能体环境中收敛至 ε-纳什均衡;同时揭示了 AIXI 的不可计算性,以及基于先验的最优性准则的局限性。

ABSTRACT

Reinforcement learning (RL) problems are often phrased in terms of Markov decision processes (MDPs). In this thesis we go beyond MDPs and consider RL in environments that are non-Markovian, non-ergodic and only partially observable. Our focus is not on practical algorithms, but rather on the fundamental underlying problems: How do we balance exploration and exploitation? How do we explore optimally? When is an agent optimal? We follow the nonparametric realizable paradigm. We establish negative results on Bayesian RL agents, in particular AIXI. We show that unlucky or adversarial choices of the prior cause the agent to misbehave drastically. Therefore Legg-Hutter intelligence and balanced Pareto optimality, which depend crucially on the choice of the prior, are entirely subjective. Moreover, in the class of all computable environments every policy is Pareto optimal. This undermines all existing optimality properties for AIXI. However, there are Bayesian approaches to general RL that satisfy objective optimality guarantees: We prove that Thompson sampling is asymptotically optimal in stochastic environments in the sense that its value converges to the value of the optimal policy. We connect asymptotic optimality to regret given a recoverability assumption on the environment that allows the agent to recover from mistakes. Hence Thompson sampling achieves sublinear regret in these environments. Our results culminate in a formal solution to the grain of truth problem: A Bayesian agent acting in a multi-agent environment learns to predict the other agents' policies if its prior assigns positive probability to them (the prior contains a grain of truth). We construct a large but limit computable class containing a grain of truth and show that agents based on Thompson sampling over this class converge to play Nash equilibria in arbitrary unknown computable multi-agent environments.

研究动机与目标

  • 解决超出马尔可夫决策过程的通用强化学习中的基本挑战,特别是在非马尔可夫性、非遍历性和部分可观察环境中。
  • 研究贝叶斯强化学习智能体(尤其是 AIXI)的极限,揭示基于先验的最优性准则(如 Legg-Hutter 智能)的主观性。
  • 在通用强化学习设置中为贝叶斯智能体建立客观最优性保证,特别是通过 Thompson 抽样。
  • 通过构建一个支持收敛至 ε-纳什均衡的极限可计算先验类,解决多智能体环境中的‘真相颗粒’问题。
  • 使用算术层级分析通用强化学习智能体(包括 AIXI 和知识寻求型智能体)的可计算性与不可计算性。

提出的方法

  • 采用非参数可实现范式:假设数据来自已知可数候选类中的未知源。
  • 应用 Carathéodory 的延拓定理,从定义在有限字符串上的函数 q 构造无限序列上的概率测度。
  • 通过将价值收敛至最优策略与可恢复性下的次线性遗憾相联系,证明 Thompson 抽样在随机环境中渐近最优。
  • 表明 AIXI 不是极限可计算的,并位于算术层级的高层级,其不可计算性通过上下界证明。
  • 构建了一个大而极限可计算的环境类,其中包含真相颗粒,从而在多智能体设置中实现收敛至 ε-纳什均衡。
  • 使用测度论工具,包括 σ-次可加性和乘积拓扑中的紧致性,证明诱导概率测度的存在性与唯一性。

实验结果

研究问题

  • RQ1Thompson 抽样能否在随机、非马尔可夫性和部分可观察环境中实现渐近最优?
  • RQ2AIXI 的最优性在多大程度上依赖于先验选择?这种依赖能否被客观证明?
  • RQ3是否存在一个贝叶斯强化学习智能体,可在任意未知可计算的多智能体环境中收敛至 ε-纳什均衡?
  • RQ4AIXI 及相关通用智能体在算术层级中的确切不可计算层级为何?
  • RQ5能否构建一个贝叶斯智能体,使其在通用强化学习中满足独立于主观先验选择的客观最优性保证?

主要发现

  • Thompson 抽样在随机环境中渐近最优,其价值收敛至最优策略的价值。
  • 在可恢复性假设下,Thompson 抽样在随机环境中实现次线性遗憾。
  • AIXI 不是极限可计算的,并位于算术层级的高层级,因此本质上不可计算。
  • 存在 AIXI 的极限可计算 ε-最优近似,为实际实现提供了可计算替代方案。
  • 在所有可计算环境的类中,每个策略都是帕累托最优的,这削弱了基于先验的最优性保证(如 AIXI 的保证)。
  • 使用包含真相颗粒的极限可计算先验类进行 Thompson 抽样的贝叶斯智能体,在任意未知可计算的多智能体环境中收敛至 ε-纳什均衡。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。