Skip to main content
QUICK REVIEW

[Paper Review] Some essential skills and their combination in an architecture for a cognitive and interactive robot

Sandra Devin, Grégoire Milliez|arXiv (Cornell University)|Mar 2, 2016
AI-based Problem Solving and Planning23 references3 citations
TL;DR

This paper proposes a cognitive architecture for human-robot collaboration that integrates essential joint action skills—such as shared intention, common ground, joint attention, action observation, co-representation, and coordination—into a multi-level control framework. The architecture enables robots to perceive human partners, infer intentions, negotiate plans, and coordinate actions through integrated perception, mental state modeling, and situated communication, advancing natural and safe human-robot interaction.

ABSTRACT

The topic of joint actions has been deeply studied in the context of Human-Human interaction in order to understand how humans cooperate. Creating autonomous robots that collaborate with humans is a complex problem, where it is relevant to apply what has been learned in the context of Human-Human interaction. The question is what skills to implement and how to integrate them in order to build a cognitive architecture, allowing a robot to collaborate efficiently and naturally with humans. In this paper, we first list a set of skills that we consider essential for Joint Action, then we analyze the problem from the robot's point of view and discuss how they can be instantiated in human-robot scenarios. Finally, we open the discussion on how to integrate such skills into a cognitive architecture for human-robot collaborative problem solving and task achievement.

Motivation & Objective

  • To identify and formalize essential skills required for effective human-robot joint action based on insights from psychology and philosophy.
  • To adapt human joint action skills—such as shared intention, common ground, and theory of mind—for robotic implementation in human-robot interaction (HRI).
  • To design a multi-level cognitive architecture that integrates perception, mental state modeling, communication, and plan execution for fluid, natural collaboration.
  • To ensure robot actions are safe, acceptable, and coordinated with human partners through real-time monitoring and adaptive response mechanisms.

Proposed method

  • Uses a three-level control architecture: Distal (goal and plan commitment), Shared Proximal (high-level plan execution), and Coupled Motor (spatiotemporal coordination).
  • Employs a Situation Assessment component to generate semantic world facts using sensor data, including relational positions, affordances, and agent postures.
  • Introduces a Mental State Management component to maintain separate, consistent representations of robot and human beliefs, goals, and capabilities.
  • Incorporates a Communication for Joint Actions component enabling situated dialogue, perspective-taking, and interpretation of verbal/non-verbal signals.
  • Uses an Intention Prediction component to infer human intentions from observed behavior and mental state estimates, guiding robot assistance decisions.
  • Employs a Shared Plan Elaboration component with a human-aware task planner (e.g., HATP) to negotiate and generate socially aware plans, and a Shared Plan Achievement component to coordinate execution with real-time monitoring of human actions.

Experimental results

Research questions

  • RQ1Which cognitive and social skills are essential for successful human-robot joint action, based on human-human collaboration?
  • RQ2How can human joint action skills—such as joint attention, action observation, and theory of mind—be adapted and implemented in robotic systems?
  • RQ3How can a cognitive architecture integrate perception, mental state modeling, communication, and plan execution to enable fluid, safe, and acceptable human-robot collaboration?
  • RQ4What role does common ground play in enabling mutual understanding during human-robot interaction, and how can it be dynamically maintained?
  • RQ5How can robots coordinate actions with humans through both planned and emergent coordination mechanisms in real time?

Key findings

  • The proposed cognitive architecture integrates essential joint action skills—shared intention, common ground, joint attention, action observation, co-representation, and coordination—into a unified framework for human-robot interaction.
  • Semantic world representation based on spatial and temporal reasoning enables the robot to interpret human references (e.g., 'on the table') and generate contextually appropriate responses.
  • Mental state modeling allows the robot to track human knowledge, goals, and capabilities, supporting perspective-taking and intention inference during collaboration.
  • The architecture supports dynamic plan negotiation and shared plan execution through a combination of high-level planning and real-time coordination, ensuring safe and legible robot actions.
  • The integration of communication components enables the robot to produce and interpret human-like verbal and non-verbal signals, enhancing mutual understanding and interaction fluency.
  • Despite progress, the paper acknowledges that a fully realized, robust cognitive architecture for fluid human-robot joint action remains an open challenge requiring further development.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.