[논문 리뷰] Learning 3D Human Shape and Pose from Dense Body Parts
이 논문은 단일 이미지에서 3D 인간 자세 및 형태 추정을 향상시키기 위해 고밀도 신체 부위 대응 관계를 활용하는 DaNet이라는 분해 및 집합 네트워크를 제안한다. 국소 및 전역 스트림을 사용하고 위치 보조 회전 보정 및 PartDrop 정규화를 통해 Human3.6M, UP3D, COCO, 3DPW에서 최신 기술 수준의 성능을 달성하며, 자세 이탈을 크게 줄이고 재구성 정확도를 향상시킨다.
Reconstructing 3D human shape and pose from monocular images is challenging despite the promising results achieved by the most recent learning-based methods. The commonly occurred misalignment comes from the facts that the mapping from images to the model space is highly non-linear and the rotation-based pose representation of body models is prone to result in the drift of joint positions. In this work, we investigate learning 3D human shape and pose from dense correspondences of body parts and propose a Decompose-and-aggregate Network (DaNet) to address these issues. DaNet adopts the dense correspondence maps, which densely build a bridge between 2D pixels and 3D vertices, as intermediate representations to facilitate the learning of 2D-to-3D mapping. The prediction modules of DaNet are decomposed into one global stream and multiple local streams to enable global and fine-grained perceptions for the shape and pose predictions, respectively. Messages from local streams are further aggregated to enhance the robust prediction of the rotation-based poses, where a position-aided rotation feature refinement strategy is proposed to exploit spatial relationships between body joints. Moreover, a Part-based Dropout (PartDrop) strategy is introduced to drop out dense information from intermediate representations during training, encouraging the network to focus on more complementary body parts as well as neighboring position features. The efficacy of the proposed method is validated on both indoor and real-world datasets including Human3.6M, UP3D, COCO, and 3DPW, showing that our method could significantly improve the reconstruction performance in comparison with previous state-of-the-art methods. Our code is publicly available at https://hongwenzhang.github.io/dense2mesh .
연구 동기 및 목표
- 단일 이미지 기반 3D 인간 재구성에서 비선형 2D에서 3D로의 매핑 및 자세 이탈 문제를 해결하기 위해.
- 2D 이미지와 3D 정점 간의 고밀도 대응 맵을 활용하여 3D 인간 자세 및 형태 추정을 향상시키기 위해.
- 분해된 전역 및 국소 인식 스트림을 통해 예측 강인성을 향상시키기 위해.
- 공간적으로 인지된 특징 보정을 통해 기반 회전 자세 이탈을 줄이기 위해.
- PartDrop를 통해 보완적인 신체 부위에 집중하도록 유도하여 일반화 능력을 향상시키기 위해.
제안 방법
- 네트워크는 2D 픽셀과 3D 메쉬 정점 간의 다리를 놓기 위해 고밀도 대응 맵을 중간 표현으로 사용한다.
- 전역 스트림과 다수의 국소 스트림을 갖춘 이중 스트림 아키텍처를 사용하여 공동 형태 및 자세 예측을 수행한다.
- 국소 스트림 특징은 자세 정확도 향상을 위해 위치 보조 회전 특징 보정 모듈을 통해 집계된다.
- 학습 중에 고밀도 특징을 무작위로 마스킹하는 파트 기반 드롭아웃 (PartDrop) 전략을 통해 결함이나 모호한 신체 부위에 대한 강인성을 유도한다.
- 3D 메쉬 정점과 관절 위치에 대한 지도를 사용하여 다중 시점 및 단일 이미지 데이터셋에서 엔드 투 엔드로 모델을 훈련시킨다.
- 메서드는 기반 회전 신체 모델 표현을 사용하며, 관절 위치 이탈을 최소화하기 위해 보정을 수행한다.
실험 결과
연구 질문
- RQ1고밀도 대응 맵은 단일 이미지에서 3D 인간 자세 및 형태 추정의 정확도와 강인성을 향상시킬 수 있는가?
- RQ2예측을 전역 및 국소 스트림으로 분해하면 3D 재구성 성능에 어떤 영향을 미치는가?
- RQ3위치 보조 특징 보정은 기반 회전 자세 표현에서의 이탈을 줄일 수 있는가?
- RQ4PartDrop은 보완적인 신체 부위에 집중하도록 유도하여 일반화 능력을 향상시키는가?
- RQ5제안된 방법은 다양한 벤치마크에서 최신 기술 수준의 접근 방식을 얼마나 뛰어나게 성능을 내는가?
주요 결과
- DaNet은 Human3.6M에서 최신 기술 수준의 성능을 달성하였으며, 이전 최신 기술 수준 방법 대비 MPJPE에서 3.2 mm 향상되었다.
- 3DPW 데이터셋에서 메서드는 MPJPE를 58.7 mm로 줄였으며, 이는 이전 접근 방식보다 뚜렷한 우월성을 보였다.
- PartDrop 전략은 특히 가림이 발생하는 어려운 실생활 이미지에서 일반화 능력을 향상시켰다.
- 위치 보조 회전 보정은 복잡한 자세에서 관절 위치 이탈을 특히 줄였다.
- 모델는 실내 및 실생활 데이터셋인 COCO와 UP3D를 포함하여 다양한 환경에서 잘 일반화되었다.
- 부분적 가림 및 저품질 입력 이미지에 대해 뛰어난 강인성을 보였다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.