[Paper Review] Intrinsically Motivated Goal Exploration Processes with Automatic Curriculum Learning
The paper formalizes Intrinsically Motivated Goal Exploration Processes (IMGEP) and introduces a Modular Population-Based IMGEP Architecture (AMB) with automatic curriculum learning, validated across 2D, Minecraft, and real humanoid robot experiments to discover diverse skills and stepping-stone capabilities.
Intrinsically motivated spontaneous exploration is a key enabler of autonomous developmental learning in human children. It enables the discovery of skill repertoires through autotelic learning, i.e. the self-generation, self-selection, self-ordering and self-experimentation of learning goals. We present an algorithmic approach called Intrinsically Motivated Goal Exploration Processes (IMGEP) to enable similar properties of autonomous learning in machines. The IMGEP architecture relies on several principles: 1) self-generation of goals, generalized as parameterized fitness functions; 2) selection of goals based on intrinsic rewards; 3) exploration with incremental goal-parameterized policy search and exploitation with a batch learning algorithm; 4) systematic reuse of information acquired when targeting a goal for improving towards other goals. We present a particularly efficient form of IMGEP, called AMB, that uses a population-based policy and an object-centered spatio-temporal modularity. We provide several implementations of this architecture and demonstrate their ability to automatically generate a learning curriculum within several experimental setups. One of these experiments includes a real humanoid robot exploring multiple spaces of goals with several hundred continuous dimensions and with distractors. While no particular target goal is provided to these autotelic agents, this curriculum allows the discovery of diverse skills that act as stepping stones for learning more complex skills, e.g. nested tool use.
Motivation & Objective
- Formalize Intrinsically Motivated Goal Exploration Processes (IMGEP) as a general framework for self-generated goals and curricula.
- Introduce AMB, a modular population-based IMGEP architecture with object-centered goal spaces and stepping-stone preserving mutations.
- Demonstrate automatic curriculum learning and efficient skill discovery through diverse experiments including robotics and real humanoid robots.
- Show that self-organized exploration yields diverse skills and enables complex capabilities via stepping-stones.
- Compare modular IMGEP variants with baselines to assess sample efficiency and curriculum quality.
Proposed method
- Define goals as parameterized fitness functions over full trajectories, enabling abstract goal spaces and diverse objective forms.
- Propose IMGEP architecture with parallel exploration and exploitation loops and data reuse across goals.
- Implement intrinsic rewards based on competence progress to guide goal selection and learning focus.
- Develop Modular Population-Based IMGEP (AMB): object-centered modular goal spaces, population-based policies, and SSPMutation to preserve stepping-stones during mutations.
- Use learning-progress driven goal sampling (via a goal space policy) and a fast memory-based meta-policy for exploration; enable asynchronous offline/batch training for exploitation.
- Provide variants such as Active Model Babbling (AMB) and Random Model Babbling (RMB) to study the impact of goal-space sampling and mutation strategies.
Experimental results
Research questions
- RQ1Can intrinsically motivated exploration autonomously generate a learning curriculum across open-ended goal spaces?
- RQ2Does modular, object-centered goal construction improve sample efficiency and diversity of discovered skills?
- RQ3How do stepping-stone preserving mutations affect tool-use and complex skill acquisition?
- RQ4How does AMB compare to baseline RMB in terms of exploration efficiency and skill diversity?
- RQ5To what extent can automatic curriculum learning transfer to real robotic setups with high-dimensional sensory inputs?
Key findings
- Intrinsic rewards based on learning progress effectively bias exploration toward goals with informative competence improvements.
- Modular, object-centered goal spaces enable structured exploration and facilitate reuse of knowledge across goals, improving skill discovery.
- Stepping-Stone Preserving Mutations (SSPMutation) help maintain progress on tool-use tasks by aligning mutations with task structure, aiding exploration around stepping-stones.
- AMB variants driven by learning-progress sampling show improved sample efficiency and diversity of behaviors compared to baselines, including in real humanoid-robot experiments.
- Autonomously generated curricula enable discovery of diverse skills and stepping-stones (e.g., nested tool use) without explicit target goals or hand-crafted curricula.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.