Skip to main content
QUICK REVIEW

[Paper Review] Explicit World Models for Reliable Human-Robot Collaboration

Kenneth Kwok, Basura Fernando|arXiv (Cornell University)|Jan 5, 2026
Social Robot Interaction and HRI0 citations
TL;DR

The paper argues for building and updating explicit world models as common ground between humans and robots to enable reliable, context-aware human-robot collaboration, rather than relying on opaque end-to-end models.

ABSTRACT

This paper addresses the topic of robustness under sensing noise, ambiguous instructions, and human-robot interaction. We take a radically different tack to the issue of reliable embodied AI: instead of focusing on formal verification methods aimed at achieving model predictability and robustness, we emphasise the dynamic, ambiguous and subjective nature of human-robot interactions that requires embodied AI systems to perceive, interpret, and respond to human intentions in a manner that is consistent, comprehensible and aligned with human expectations. We argue that when embodied agents operate in human environments that are inherently social, multimodal, and fluid, reliability is contextually determined and only has meaning in relation to the goals and expectations of humans involved in the interaction. This calls for a fundamentally different approach to achieving reliable embodied AI that is centred on building and updating an accessible "explicit world model" representing the common ground between human and AI, that is used to align robot behaviours with human expectations.

Motivation & Objective

  • Motivate a shift from end-to-end black-box control to reliable collaboration built on explicit world models.
  • Highlight how common ground and multimodal grounding support interpretability and alignment with human goals.
  • Survey existing work on perceptual grounding, joint attention, and neuro-symbolic architectures to motivate explicit world modelling.

Proposed method

  • Discuss symbolic and neuro-symbolic world models as foundations for explicit representations of environment, states, and actions.
  • Explain how explicit world models can serve as common ground to resolve ambiguity and subjective interpretations in HRC.
  • Review prior work on perceptual grounding, joint attention, multimodal cues, and legible robot behavior to support the approach.
  • Propose light-weight, real-time updating of explicit world models to capture social, multimodal dynamics in human-robot interaction.

Experimental results

Research questions

  • RQ1How can explicit world models be constructed and maintained to serve as common ground in human-robot collaboration?
  • RQ2What role do multimodal cues (gaze, gestures, prosody) and joint attention play in building reliable explicit world models?
  • RQ3Can neuro-symbolic architectures provide interpretable, verifiable reasoning for HRC tasks within explicit world models?

Key findings

  • Explicit world models offer a path to reliability by anchoring robot behavior to a shared interpretation of states and human intentions.
  • Explicit, interpretable representations can resolve ambiguity and subjectivity better than opaque end-to-end models in dynamic human environments.
  • A synthesis of symbolic, neuro-symbolic, and multimodal grounding literature supports building communicable common ground for HRC.
  • Real-time, lightweight world models are needed to capture social and multimodal dynamics without sacrificing responsiveness.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.