Skip to main content
QUICK REVIEW

[论文解读] Hierarchical principles of embodied reinforcement learning: A review

Manfred Eppe, Christian Gumbsch|arXiv (Cornell University)|Dec 18, 2020
Child and Animal Learning Development参考文献 128被引用 5
一句话总结

本文通过将动物和人类问题解决中的认知机制(如组合抽象、前向模型、好奇心和心理模拟)与计算型分层强化学习(HRL)方法相联系,回顾了分层强化学习(HRL)。研究指出,尽管关键组件在孤立状态下已存在,但其整合仍缺失,因此提出将这些机制整合是实现人工智能代理具备零样本、动物级问题解决能力的关键。

ABSTRACT

Cognitive Psychology and related disciplines have identified several critical mechanisms that enable intelligent biological agents to learn to solve complex problems. There exists pressing evidence that the cognitive mechanisms that enable problem-solving skills in these species build on hierarchical mental representations. Among the most promising computational approaches to provide comparable learning-based problem-solving abilities for artificial agents and robots is hierarchical reinforcement learning. However, so far the existing computational approaches have not been able to equip artificial agents with problem-solving abilities that are comparable to intelligent animals, including human and non-human primates, crows, or octopuses. Here, we first survey the literature in Cognitive Psychology, and related disciplines, and find that many important mental mechanisms involve compositional abstraction, curiosity, and forward models. We then relate these insights with contemporary hierarchical reinforcement learning methods, and identify the key machine intelligence approaches that realise these mechanisms. As our main result, we show that all important cognitive mechanisms have been implemented independently in isolated computational architectures, and there is simply a lack of approaches that integrate them appropriately. We expect our results to guide the development of more sophisticated cognitively inspired hierarchical methods, so that future artificial agents achieve a problem-solving performance on the level of intelligent animals.

研究动机与目标

  • 识别使乌鸦、灵长类动物和章鱼等智能动物实现零样本问题解决的认知机制。
  • 分析当前分层强化学习(HRL)方法是否实现了这些认知机制。
  • 识别现有HRL方法中的关键缺陷,尤其是缺乏集成的前向模型、组合抽象与内在动机。
  • 提出构建动物级智能代理的工具已存在,但需实现整体性整合。
  • 引导未来研究朝向开发更具认知启发性的、分层的强化学习系统。

提出的方法

  • 开展自上而下的认知心理学与动物认知文献调查,提取用于分层问题解决的核心机制。
  • 将这些认知机制——组合抽象、前向模型、内在动机(好奇心/多样性)和心理模拟——映射到计算型HRL框架中。
  • 系统性回顾37篇近期强化学习综述文章(2015–2020年)及117种HRL架构,评估这些机制的实现情况。
  • 识别并分析现有HRL系统中关键机制(如前向模型、组合表示)的孤立实现。
  • 使用表格分析(表1)比较HRL方法在以下维度的表现:前向模型、状态抽象、动作抽象、内在动机与心理模拟。
  • 指出大多数HRL方法为无模型方法,缺乏组合抽象,并且未整合多种认知机制。

实验结果

研究问题

  • RQ1乌鸦和灵长类动物等智能动物实现零样本问题解决的认知机制是什么?
  • RQ2当前的分层强化学习方法在多大程度上实现了这些认知机制?
  • RQ3为何现有HRL方法尽管理论前景良好,却仍无法实现动物级的问题解决性能?
  • RQ4哪些关键组件——如前向模型、组合抽象或内在动机——在当前HRL系统中缺失或代表性不足?
  • RQ5将孤立存在的现有机制进行整合,能否使人工智能代理具备人类与动物级的少样本问题解决能力?

主要发现

  • 在37篇2015–2020年间的近期强化学习综述文章中,仅有2篇(占5.4%)在摘要中使用了‘hierarchic’一词,表明主流文献对分层方法的关注度极低。
  • Sutton与Barto的《强化学习》教科书第二版仅用2页提及分层结构,且主要局限于过时的选项框架,凸显主流强化学习理论中对分层结构的整合程度有限。
  • 大多数HRL方法为无模型方法,未整合预测性前向模型,尽管有强有力的认知证据表明其在规划与抽象中的关键作用。
  • 组合抽象极少被支持:仅有少数方法使用自然语言或符号逻辑表示动作,且无一将这些表示与感知运动经验紧密对齐。
  • 状态抽象显著研究不足;大多数分层演员-critic方法在所有层级上使用相同的状态表示,未实现抽象。
  • 内在动机机制——尤其是好奇心与多样性——在HRL中代表性不足,尽管其在驱动生物体认知探索方面的作用已有充分记录。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。