[Paper Review] A Dataset of Relighted 3D Interacting Hands
This paper introduces Re:InterHand, a large-scale dataset of relighted 3D interacting hands that combines realistic, diverse image appearances with accurate 3D ground-truth poses. By applying a state-of-the-art hand relighting network to precisely tracked 3D hand poses from a multi-camera studio, the dataset achieves high realism and diversity in appearance while maintaining strong 3D pose accuracy, outperforming existing datasets in both realism and annotation quality.
The two-hand interaction is one of the most challenging signals to analyze due to the self-similarity, complicated articulations, and occlusions of hands. Although several datasets have been proposed for the two-hand interaction analysis, all of them do not achieve 1) diverse and realistic image appearances and 2) diverse and large-scale groundtruth (GT) 3D poses at the same time. In this work, we propose Re:InterHand, a dataset of relighted 3D interacting hands that achieve the two goals. To this end, we employ a state-of-the-art hand relighting network with our accurately tracked two-hand 3D poses. We compare our Re:InterHand with existing 3D interacting hands datasets and show the benefit of it. Our Re:InterHand is available in https://mks0601.github.io/ReInterHand/.
Motivation & Objective
- Address the lack of datasets that simultaneously provide diverse, realistic image appearances and large-scale, accurate 3D hand pose annotations for two-hand interactions.
- Overcome limitations of existing datasets—lab-based (monotonous appearance), natural (limited scale and accuracy), and composited (lighting inconsistency)—by combining the strengths of each approach.
- Enable more robust 3D hand pose estimation in real-world (in-the-wild) scenarios by generating images with realistic lighting and reflections using a learned relighting network.
- Provide a benchmark for evaluating 3D interacting hand recovery methods under realistic visual conditions, including egocentric and third-person viewpoints.
- Facilitate research in 3D hand pose estimation by offering a large-scale, high-quality dataset with diverse hand poses and photorealistic rendering.
Proposed method
- Leverage a multi-camera studio setup to capture accurate 3D hand poses with 100 synchronized cameras, ensuring high-fidelity ground-truth 3D joint and MANO mesh annotations.
- Apply a state-of-the-art hand relighting network trained on single-hand data to render 3D hand meshes with realistic lighting, using high-resolution environment maps for diverse illumination conditions.
- Generate 1.5 million images by rendering 3D hand meshes with varying poses, lighting, and backgrounds, preserving realistic skin tones and reflections.
- Use environment maps to simulate real-world lighting variations, enhancing appearance diversity while maintaining consistency between hand geometry and lighting.
- Apply post-processing techniques such as AdaIN to improve color harmony between foreground hands and background, though with limited success due to reflection inaccuracies.
- Ensure dataset quality by excluding identifiable information and using consent forms, preserving privacy while maintaining data utility.
Experimental results
Research questions
- RQ1Can a 3D hand dataset combine realistic, diverse image appearances with large-scale, accurate 3D pose annotations to improve generalization in real-world settings?
- RQ2How does relighting 3D hand meshes with a learned network affect the performance of 3D hand pose estimation models on real-world and in-the-wild data?
- RQ3To what extent does the proposed dataset outperform existing datasets in terms of realism, pose diversity, and 3D accuracy?
- RQ4Does training on Re:InterHand improve generalization to egocentric and third-person viewing conditions compared to existing benchmarks?
- RQ5What are the limitations of current relighting networks when applied to complex, interacting two-hand scenarios?
Key findings
- Re:InterHand achieves a 20.07 RRVE on its own test split when fine-tuned on the dataset, significantly outperforming the baseline InterWild model (28.89 RRVE) on egocentric views.
- The relighting stage reduces error on HIC (a real-world dataset) from 67.11 to 52.91 RRVE, demonstrating improved generalization to natural image appearances.
- The dataset reduces error on InterHand2.6M by 2.4% (from 19.74 to 19.40 RRVE) when fine-tuned on Re:InterHand, showing improved robustness even on lab-based data.
- The relighting process preserves skin tones and realistic lighting, outperforming simple composition methods that suffer from lighting inconsistency between foreground and background.
- Despite challenges, the relighting network generalizes reasonably well to two-hand interactions, though occasional artifacts occur due to domain shift from single-hand to two-hand data.
- The dataset enables state-of-the-art performance on both third-person and egocentric benchmarks, validating its utility for real-world 3D hand pose estimation.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.