Skip to main content
QUICK REVIEW

[论文解读] A Definition of Open-Ended Learning Problems for Goal-Conditioned Agents

Olivier Sigaud, Gianluca Baldassarre|arXiv (Cornell University)|Nov 1, 2023
Reinforcement Learning in Robotics参考文献 58被引用 4
一句话总结

本文通过隔离一个核心特性——从观察者视角出发,持续、无限时域地生成新颖的目标、选项或奖励函数——为目标条件智能体提出了开放性学习(OEL)的形式化定义。它引入了开放性目标条件强化学习(GCRL)问题,并建立了一个评估框架,以衡量智能体随时间发现并掌握新目标的能力,从而将OEL与持续学习、终身学习等类似概念区分开来。

ABSTRACT

A lot of recent machine learning research papers have ``open-ended learning'' in their title. But very few of them attempt to define what they mean when using the term. Even worse, when looking more closely there seems to be no consensus on what distinguishes open-ended learning from related concepts such as continual learning, lifelong learning or autotelic learning. In this paper, we contribute to fixing this situation. After illustrating the genealogy of the concept and more recent perspectives about what it truly means, we outline that open-ended learning is generally conceived as a composite notion encompassing a set of diverse properties. In contrast with previous approaches, we propose to isolate a key elementary property of open-ended processes, which is to produce elements from time to time (e.g., observations, options, reward functions, and goals), over an infinite horizon, that are considered novel from an observer's perspective. From there, we build the notion of open-ended learning problems and focus in particular on the subset of open-ended goal-conditioned reinforcement learning problems in which agents can learn a growing repertoire of goal-driven skills. Finally, we highlight the work that remains to be performed to fill the gap between our elementary definition and the more involved notions of open-ended learning that developmental AI researchers may have in mind.

研究动机与目标

  • 为机器学习领域中开放性学习(OEL)文献日益增长但缺乏共识与形式化定义的问题提供解决。
  • 隔离OEL的一个核心、基本属性:从观察者视角出发,在无限时间范围内持续生成新颖元素(如目标、选项)的能力。
  • 在目标条件强化学习(GCRL)的语境下,具体定义开放性学习问题。
  • 通过一个系统性框架,将OEL与持续学习、终身学习及自激励学习等类似概念区分开来。
  • 识别出捕捉技能习得发展轨迹的研究方向,包括表征重述与抽象化等。

提出的方法

  • 本文通过一个核心属性定义OEL:智能体在无限时间范围内持续生成新颖元素(如目标、选项、奖励函数),其中新颖性由外部观察者进行评估。
  • 引入开放性GCRL问题,其中智能体的学习分为两个阶段:第一阶段为内在阶段,无外部指导;第二阶段为外在阶段,从环境的任务空间中随机抽取任务。
  • 以外在阶段的表现作为内在阶段知识获取能力的代理指标,衡量智能体对未见任务的泛化能力。
  • 提出通过追踪随时间推移的新目标发现速率来评估OEL智能体,若该速率趋于收敛,则表明非OEL行为。
  • 建议在外在阶段使用统一的任务集,以实现对不同内在课程的智能体进行公平比较,从而实现对OEL智能体性能的基于表现的比较。
  • 该框架强调,必须将核心新颖性属性与额外能力(如技能抽象、表征重述、创造力)相结合,以反映人类的发展性学习。

实验结果

研究问题

  • RQ1定义自主智能体中开放性学习的根本性、基本属性是什么?
  • RQ2如何在目标条件强化学习中正式定义开放性学习问题?
  • RQ3开放性学习与持续学习、终身学习或自激励学习等概念在何种意义上存在差异?
  • RQ4当OEL智能体采用不同内在课程时,如何公平地评估其性能?
  • RQ5除了新颖性生成之外,还需哪些额外能力才能捕捉OEL智能体中技能习得的发展轨迹?

主要发现

  • 开放性学习的核心属性是从观察者视角出发,持续、无限时域地生成新颖元素(如目标、选项、奖励函数)。
  • 本文建立了开放性目标条件强化学习(GCRL)问题的形式化框架,将其与标准强化学习设置区分开来。
  • 若智能体随时间推移无法发现新目标(表现为已见目标数量趋于收敛),则不构成开放性学习智能体。
  • 随时间推移,新目标发现速率呈对数或线性增长,表明存在开放性学习行为,且增长率上升是最强的指示信号。
  • 在外在阶段使用标准化任务集进行随机任务测试,可作为衡量OEL过程在内在阶段知识获取能力的有效代理指标。
  • 该框架强调,必须将额外的发展机制(如表征重述、抽象化、创造力)整合进OEL智能体中,以反映类人技能习得过程。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。