Skip to main content
QUICK REVIEW

[Paper Review] A Dataset of Relighted 3D Interacting Hands

Gyeongsik Moon, Shunsuke Saito|arXiv (Cornell University)|Oct 26, 2023
Hand Gesture Recognition Systems4 citations
TL;DR

This paper introduces Re:InterHand, a large-scale dataset of relighted 3D interacting hands that combines realistic, diverse image appearances with accurate 3D ground-truth poses. By applying a state-of-the-art hand relighting network to precisely tracked 3D hand poses from a multi-camera studio, the dataset achieves high realism and diversity in appearance while maintaining strong 3D pose accuracy, outperforming existing datasets in both realism and annotation quality.

ABSTRACT

The two-hand interaction is one of the most challenging signals to analyze due to the self-similarity, complicated articulations, and occlusions of hands. Although several datasets have been proposed for the two-hand interaction analysis, all of them do not achieve 1) diverse and realistic image appearances and 2) diverse and large-scale groundtruth (GT) 3D poses at the same time. In this work, we propose Re:InterHand, a dataset of relighted 3D interacting hands that achieve the two goals. To this end, we employ a state-of-the-art hand relighting network with our accurately tracked two-hand 3D poses. We compare our Re:InterHand with existing 3D interacting hands datasets and show the benefit of it. Our Re:InterHand is available in https://mks0601.github.io/ReInterHand/.

Motivation & Objective

  • Address the lack of datasets that simultaneously provide diverse, realistic image appearances and large-scale, accurate 3D hand pose annotations for two-hand interactions.
  • Overcome limitations of existing datasets—lab-based (monotonous appearance), natural (limited scale and accuracy), and composited (lighting inconsistency)—by combining the strengths of each approach.
  • Enable more robust 3D hand pose estimation in real-world (in-the-wild) scenarios by generating images with realistic lighting and reflections using a learned relighting network.
  • Provide a benchmark for evaluating 3D interacting hand recovery methods under realistic visual conditions, including egocentric and third-person viewpoints.
  • Facilitate research in 3D hand pose estimation by offering a large-scale, high-quality dataset with diverse hand poses and photorealistic rendering.

Proposed method

  • Leverage a multi-camera studio setup to capture accurate 3D hand poses with 100 synchronized cameras, ensuring high-fidelity ground-truth 3D joint and MANO mesh annotations.
  • Apply a state-of-the-art hand relighting network trained on single-hand data to render 3D hand meshes with realistic lighting, using high-resolution environment maps for diverse illumination conditions.
  • Generate 1.5 million images by rendering 3D hand meshes with varying poses, lighting, and backgrounds, preserving realistic skin tones and reflections.
  • Use environment maps to simulate real-world lighting variations, enhancing appearance diversity while maintaining consistency between hand geometry and lighting.
  • Apply post-processing techniques such as AdaIN to improve color harmony between foreground hands and background, though with limited success due to reflection inaccuracies.
  • Ensure dataset quality by excluding identifiable information and using consent forms, preserving privacy while maintaining data utility.

Experimental results

Research questions

  • RQ1Can a 3D hand dataset combine realistic, diverse image appearances with large-scale, accurate 3D pose annotations to improve generalization in real-world settings?
  • RQ2How does relighting 3D hand meshes with a learned network affect the performance of 3D hand pose estimation models on real-world and in-the-wild data?
  • RQ3To what extent does the proposed dataset outperform existing datasets in terms of realism, pose diversity, and 3D accuracy?
  • RQ4Does training on Re:InterHand improve generalization to egocentric and third-person viewing conditions compared to existing benchmarks?
  • RQ5What are the limitations of current relighting networks when applied to complex, interacting two-hand scenarios?

Key findings

  • Re:InterHand achieves a 20.07 RRVE on its own test split when fine-tuned on the dataset, significantly outperforming the baseline InterWild model (28.89 RRVE) on egocentric views.
  • The relighting stage reduces error on HIC (a real-world dataset) from 67.11 to 52.91 RRVE, demonstrating improved generalization to natural image appearances.
  • The dataset reduces error on InterHand2.6M by 2.4% (from 19.74 to 19.40 RRVE) when fine-tuned on Re:InterHand, showing improved robustness even on lab-based data.
  • The relighting process preserves skin tones and realistic lighting, outperforming simple composition methods that suffer from lighting inconsistency between foreground and background.
  • Despite challenges, the relighting network generalizes reasonably well to two-hand interactions, though occasional artifacts occur due to domain shift from single-hand to two-hand data.
  • The dataset enables state-of-the-art performance on both third-person and egocentric benchmarks, validating its utility for real-world 3D hand pose estimation.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.