[논문 리뷰] PVN3D: A Deep Point-wise 3D Keypoints Voting Network for 6DoF Pose Estimation
PVN3D는 단일 RGBD 이미지에서 6DoF 포즈를 추정하기 위해 인스턴스 세그먼테이션을 갖춘 3D 키포인트 보팅 네트워크를 제시하며, YCB-Video 및 LineMOD에서 기존 방법을 능가합니다.
In this work, we present a novel data-driven method for robust 6DoF object pose estimation from a single RGBD image. Unlike previous methods that directly regressing pose parameters, we tackle this challenging task with a keypoint-based approach. Specifically, we propose a deep Hough voting network to detect 3D keypoints of objects and then estimate the 6D pose parameters within a least-squares fitting manner. Our method is a natural extension of 2D-keypoint approaches that successfully work on RGB based 6DoF estimation. It allows us to fully utilize the geometric constraint of rigid objects with the extra depth information and is easy for a network to learn and optimize. Extensive experiments were conducted to demonstrate the effectiveness of 3D-keypoint detection in the 6D pose estimation task. Experimental results also show our method outperforms the state-of-the-art methods by large margins on several benchmarks. Code and video are available at https://github.com/ethnhe/PVN3D.git.
연구 동기 및 목표
- 직접 포즈 회귀가 아닌 3D 키포인트를 활용하여 RGB-D에서 견고한 6DoF 포즈 추정을 촉진한다.
제안 방법
- 선택된 3D 키포인트에 대해 포인트별 평행이동 오프셋를 학습하고 키포인트를 얻기 위해 클러스터링을 사용하는 심층 3D 키포인트 허프 보팅 네트워크를 개발한다.
실험 결과
연구 질문
- RQ13D 키포인트 기반 보팅이 6DoF 포즈 추정에서 직접 포즈 회귀 및 2D 키포인트 접근 방식보다 더 나은 성능을 보일 수 있는가?
주요 결과
- ADD-S 및 ADD-S AUC 지표에서 YCB-Video 및 LineMOD 데이터셋에서 최첨단 방법들을 능가한다.
- 실험에서 3D 키포인트 형식이 직접 포즈 회귀 및 2D 키포인트 접근 방식보다 우수하다.
- 3D 키포인트 보팅을 인스턴스 의미론적 세분화와 함께 공동 학습하면 포즈 정확도와 세그먼트 품질이 향상된다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.