[Paper Review] DeepMultiCap: Performance Capture of Multiple Characters Using Sparse Multiview Cameras
DeepMultiCap presents a novel method for high-fidelity, multi-person performance capture from sparse multi-view RGB videos without requiring pre-scanned templates. By combining a spatial attention-aware module, SMPL-based 3D priors, and a temporal fusion strategy, it achieves state-of-the-art reconstruction under severe occlusions and temporal inconsistencies, validated on a new high-quality dataset, MultiHuman.
We propose DeepMultiCap, a novel method for multi-person performance capture using sparse multi-view cameras. Our method can capture time varying surface details without the need of using pre-scanned template models. To tackle with the serious occlusion challenge for close interacting scenes, we combine a recently proposed pixel-aligned implicit function with parametric model for robust reconstruction of the invisible surface areas. An effective attention-aware module is designed to obtain the fine-grained geometry details from multi-view images, where high-fidelity results can be generated. In addition to the spatial attention method, for video inputs, we further propose a novel temporal fusion method to alleviate the noise and temporal inconsistencies for moving character reconstruction. For quantitative evaluation, we contribute a high quality multi-person dataset, MultiHuman, which consists of 150 static scenes with different levels of occlusions and ground truth 3D human models. Experimental results demonstrate the state-of-the-art performance of our method and the well generalization to real multiview video data, which outperforms the prior works by a large margin.
Motivation & Objective
- To enable high-fidelity, multi-person 3D performance capture from sparse multi-view RGB videos without pre-scanned templates.
- To address severe occlusions in close-interacting scenes by combining implicit functions with parametric SMPL models as 3D geometry proxies.
- To improve fine-grained geometric detail recovery through a spatial attention-aware module for multi-view feature fusion.
- To enhance temporal consistency in dynamic sequences using a novel SDF-based temporal fusion method.
- To provide a new, high-quality benchmark dataset, MultiHuman, for evaluating multi-person performance capture systems.
Proposed method
- A spatial attention-aware module is designed to adaptively aggregate multi-view features, emphasizing high-frequency geometric details from different viewpoints.
- The method integrates SMPL as a 3D prior to reconstruct complete human bodies in occluded regions, with a global normal map to guide attention and improve robustness.
- A novel temporal fusion strategy weights signed distance fields (SDFs) across video frames to reduce noise and improve temporal consistency in dynamic reconstructions.
- The framework uses a pixel-aligned implicit function to represent 3D geometry, enabling efficient and detailed surface reconstruction from multi-view images.
- The system is trained and evaluated using a new dataset, MultiHuman, containing 150 high-quality, multi-person, interactive scenes with varying occlusion levels.
Experimental results
Research questions
- RQ1Can a deep learning-based method achieve high-fidelity 3D reconstruction of multiple interacting humans from sparse multi-view RGB videos without pre-scanned templates?
- RQ2How can severe occlusions in multi-person performance capture be mitigated using 3D priors and attention mechanisms?
- RQ3To what extent can a spatial attention module improve fine-grained geometric detail recovery in multi-view 3D reconstruction?
- RQ4How effective is a temporal fusion strategy in reducing noise and temporal inconsistencies in dynamic 3D sequences?
- RQ5Can a new, high-quality dataset of multi-person, occluded scenes serve as a reliable benchmark for evaluating performance capture systems?
Key findings
- DeepMultiCap achieves state-of-the-art performance on multi-person performance capture, outperforming prior methods by a large margin on both synthetic and real-world data.
- The spatial attention-aware module significantly improves reconstruction quality, especially when high-frequency details like normal maps are involved.
- The integration of SMPL as a 3D proxy enables robust reconstruction of occluded body parts, with ablation studies showing a notable drop in accuracy when the SMPL prior is removed.
- The temporal fusion method reduces temporal artifacts and improves consistency in dynamic sequences, as shown in qualitative comparisons and supplementary videos.
- The MultiHuman dataset, consisting of 150 high-quality, multi-person, interactive scenes with varying occlusion levels, enables detailed and reliable evaluation of multi-person performance capture systems.
- The method demonstrates strong generalization to real-world multiview video data, producing high-fidelity results even under challenging occlusion conditions.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.