[Paper Review] A Survey on Interactive Reinforcement Learning: Design Principles and Open Challenges
This survey equips HCI researchers with technical foundations in interactive reinforcement learning (IRL) by identifying design principles, feedback integration methods, and open challenges such as user modeling, safety, and evaluation. It proposes actionable guidelines for developing effective, user-centered IRL applications across robotics, HCI, and game AI using human-in-the-loop feedback to improve learning efficiency and personalization.
Interactive reinforcement learning (RL) has been successfully used in various applications in different fields, which has also motivated HCI researchers to contribute in this area. In this paper, we survey interactive RL to empower human-computer interaction (HCI) researchers with the technical background in RL needed to design new interaction techniques and propose new applications. We elucidate the roles played by HCI researchers in interactive RL, identifying ideas and promising research directions. Furthermore, we propose generic design principles that will provide researchers with a guide to effectively implement interactive RL applications.
Motivation & Objective
- To bridge the gap between HCI researchers and interactive reinforcement learning (IRL) by providing essential technical background in RL.
- To identify and articulate the roles HCI researchers can play in IRL, particularly in designing novel interaction techniques and applications.
- To address key challenges in IRL, including user fatigue, lack of human-like oracles, and poor generalization across environments.
- To propose generic, transferable design principles for implementing effective IRL systems in real-world physical and simulated environments.
- To highlight underexplored research directions such as adaptive feedback channels, explainable AI in IRL, and safe learning mechanisms.
Proposed method
- Conducts a focused survey of IRL research from 2010–2019 in robotics, HCI, and game AI, prioritizing works with human-in-the-loop applications.
- Analyzes feedback types (e.g., reward shaping, demonstration, preference feedback) and their roles in personalization, exploration guidance, and sample efficiency.
- Proposes design principles derived from cross-domain analysis of IRL systems, emphasizing usability, adaptability, and user engagement.
- Evaluates existing evaluation practices and identifies shortcomings, such as overreliance on simple testbeds like GridWorld and lack of fast, visual feedback evaluation tools.
- Introduces the need for formal user modeling to predict user behavior and reduce fatigue, enabling proactive and adaptive interaction.
- Advocates for integrating safe RL techniques and explainable AI to ensure agent transparency, bias detection, and user trust in high-stakes applications.
Experimental results
Research questions
- RQ1How can HCI researchers effectively contribute to interactive reinforcement learning through novel interaction techniques and applications?
- RQ2What are the most effective feedback types in IRL, and how do they influence learning efficiency, personalization, and exploration?
- RQ3Why do IRL systems often fail to generalize from simple testbeds (e.g., GridWorld) to complex environments like Infinite Mario?
- RQ4How can user modeling and adaptive interaction design reduce user fatigue and improve feedback quality in IRL systems?
- RQ5What role do explainable AI and safe RL techniques play in enabling trustworthy and reliable human-agent collaboration?
Key findings
- Interactive RL is underutilized in the HCI community despite its potential for personalization and improved sample efficiency in complex tasks.
- User feedback types—such as reward shaping, demonstrations, and preference signals—can significantly enhance learning, but their effectiveness varies across environments and user expertise.
- Simulated oracles often fail to replicate human feedback behavior, leading to misleading performance predictions, highlighting the need for human-centered evaluation.
- Generalization of IRL methods from simple to complex environments remains a major challenge, as demonstrated by the failure of RS (Reward Shaping) in complex games like Infinite Mario.
- A formal user model is currently missing, limiting the ability to proactively manage user interaction frequency and reduce fatigue in long-term IRL applications.
- Explainable IRL and fast evaluation techniques (e.g., visualization of behavior and uncertainty) are underdeveloped but essential for improving feedback quality and debugging agent behavior.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.