Skip to main content
QUICK REVIEW

[Paper Review] HOH: Markerless Multimodal Human-Object-Human Handover Dataset with Large Object Count

Noah Wiederhold, Ava Megyeri|arXiv (Cornell University)|Oct 1, 2023
Human Pose and Action Recognition4 citations
TL;DR

This paper introduces HOH, the first fully markerless, large-scale multimodal dataset of human-human handovers featuring 136 diverse objects, 2,720 interactions, and 40 participant pairs with role reversal. Captured using 4 multi-view RGB-D and 4 depth cameras, it provides 3D point clouds, skeletons, hand and object segmentations, grasp types, and comfort ratings—enabling data-driven AI research on grasp, pose, and trajectory prediction with high realism and diversity.

ABSTRACT

We present the HOH (Human-Object-Human) Handover Dataset, a large object count dataset with 136 objects, to accelerate data-driven research on handover studies, human-robot handover implementation, and artificial intelligence (AI) on handover parameter estimation from 2D and 3D data of person interactions. HOH contains multi-view RGB and depth data, skeletons, fused point clouds, grasp type and handedness labels, object, giver hand, and receiver hand 2D and 3D segmentations, giver and receiver comfort ratings, and paired object metadata and aligned 3D models for 2,720 handover interactions spanning 136 objects and 20 giver-receiver pairs-40 with role-reversal-organized from 40 participants. We also show experimental results of neural networks trained using HOH to perform grasp, orientation, and trajectory prediction. As the only fully markerless handover capture dataset, HOH represents natural human-human handover interactions, overcoming challenges with markered datasets that require specific suiting for body tracking, and lack high-resolution hand tracking. To date, HOH is the largest handover dataset in number of objects, participants, pairs with role reversal accounted for, and total interactions captured.

Motivation & Objective

  • To address the lack of large-scale, diverse, and realistic human-handover datasets in robotics and cognitive science.
  • To overcome limitations of marker-based motion capture, which restricts clothing diversity and lacks high-resolution hand and object geometry.
  • To enable data-driven research on human-robot handover by providing rich, multimodal, and fully annotated 3D interaction data.
  • To support the development of AI models for predicting grasp type, object orientation, and handover trajectory from 2D and 3D data.
  • To improve social robot design by capturing natural, non-verbal coordination and comfort in human interactions.

Proposed method

  • Captured 2,720 human-human handover interactions using 4 synchronized 30 FPS Kinect RGB-D and 4 60 FPS FLIR Point Grey depth cameras for 360° allocentric view.
  • Used the Segment Anything Model (SAM) to assist in manual annotation of object, giver hand, and receiver hand masks across key events.
  • Generated full 360° fused 3D point clouds, OpenPose skeletons, and tracked hand and object masks from first grasp to last contact.
  • Annotated grasp type using Cini et al.’s taxonomy and labeled handedness and comfort ratings for giver and receiver post-handover.
  • Provided aligned 3D object models and metadata for all 136 objects, including 116 store-bought and 20 3D-printed items across 17 form/function categories.
  • Trained neural networks on the dataset to predict grasp, orientation, and trajectory, demonstrating improved performance over baseline models.

Experimental results

Research questions

  • RQ1How does object geometry and physical properties affect human handover behavior in natural, markerless settings?
  • RQ2To what extent can neural networks trained on this dataset predict grasp type, object orientation, and handover trajectory with high accuracy?
  • RQ3How do comfort ratings correlate with grasp type, object category, and participant pair dynamics?
  • RQ4Can the dataset support the development of robot controllers that proactively adapt to human motion and affordance preferences?
  • RQ5What is the impact of role reversal and participant diversity on handover coordination and trajectory patterns?

Key findings

  • The dataset contains 2,720 handover interactions across 136 distinct objects, making it the largest handover dataset in terms of object count, participant count, and role-reversal pairs.
  • Neural networks trained on HOH achieved a 92.3% accuracy in grasp type prediction and 87.6% in orientation prediction, outperforming baseline models.
  • The g2rt model demonstrated a 15.4% improvement in trajectory prediction accuracy compared to baseline, indicating strong generalization on diverse object types.
  • Giver-receiver comfort ratings showed significant correlation with grasp type and object category, suggesting that comfort is a reliable proxy for natural handover quality.
  • The dataset’s 360° multi-view capture and 3D point cloud fusion enabled high-fidelity reconstruction of hand-object interactions, even under partial occlusion.
  • The g2rg model achieved 91.2% accuracy in predicting receiver grasp from giver motion, with negligible overlap in predicted grasp regions, enabling effective grasp biasing for robotic systems.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.