Skip to main content
QUICK REVIEW

[논문 리뷰] OnePose++: Keypoint-Free One-Shot Object Pose Estimation without CAD Models

Xingyi He, Jiaming Sun|arXiv (Cornell University)|2023. 01. 18.
Robot Manipulation and Learning인용 수 29
한 줄 요약

키포인트 프리 SfM 및 희소-to-밀도 2D-3D 매칭 파이프라인으로, CAD 모델 없이 참조 이미지에서 반-밀집 객체 포인트 구름을 재구성하고 질의 이미지의 자세를 추정하며 질감이 있는 객체와 저질감 객체 모두에서 원샷 결과로 강력한 성능을 달성합니다.

ABSTRACT

We propose a new method for object pose estimation without CAD models. The previous feature-matching-based method OnePose has shown promising results under a one-shot setting which eliminates the need for CAD models or object-specific training. However, OnePose relies on detecting repeatable image keypoints and is thus prone to failure on low-textured objects. We propose a keypoint-free pose estimation pipeline to remove the need for repeatable keypoint detection. Built upon the detector-free feature matching method LoFTR, we devise a new keypoint-free SfM method to reconstruct a semi-dense point-cloud model for the object. Given a query image for object pose estimation, a 2D-3D matching network directly establishes 2D-3D correspondences between the query image and the reconstructed point-cloud model without first detecting keypoints in the image. Experiments show that the proposed pipeline outperforms existing one-shot CAD-model-free methods by a large margin and is comparable to CAD-model-based methods on LINEMOD even for low-textured objects. We also collect a new dataset composed of 80 sequences of 40 low-textured objects to facilitate future research on one-shot object pose estimation. The supplementary material, code and dataset are available on the project page: https://zju3dv.github.io/onepose_plus_plus/.

연구 동기 및 목표

  • 원샷 객체 자세 추정에서 반복 가능한 키포인트에 대한 의존성 제거.
  • 참조 뷰로부터 정확한 반-밀집 객체 포인트 구름을 재구성하기 위한 키포인트 프리 SfM 파이프라인 개발.
  • 테스트 이미지에서 효율적이고 정확한 자세 추정용 희소-대-밀도 2D-3D 매칭 네트워크 설계.
  • 표준 데이터셋에서 CAD-모델 비의 기반선 대비 향상된 성능 및 CAD 모델 기반 방법과의 경쟁력 있는 결과 시연.
  • 원샷 자세 추정을 위한 향후 연구를 촉진하는 저질감 객체 데이터셋 제공

제안 방법

  • LoFTR 스타일의 키포인트 프리 매칭을 두 단계 SfM 프레임워크에 적용: 반복 가능한 코arse 매치를 통한 거친 재구성으로 완전한 반-밀집 포인트 구름을 형성하고, 그 다음 특징 트랙과 3D 점의 미세-소수점 정확도 보정.
  • 트랙당 하나의 기준 노드를 고정하고 국소적으로 서브픽셀 보정으로 거친 트랙을 정제한 뒤 재투영 오차를 통해 3D 포인트 구름을 최적화.
  • 테스트 시점에 질의 이미지와 재구성된 포인트 구름 간에 Transformer 기반 교차-어텐션 네트워크를 이용한 거친 2D-3D 매칭을 수행하고, 그 다음 로컬 윈도우 내에서 정밀한 2D-3D 매칭을 통해 PnP에 필요한 정밀한 대응점을 얻는다.
  • 2D-3D 매칭에서 길고 범위 의존성을 모델링하기 위해 선형 어텐션을 이용한 self- 및 cross-attention을 사용하고, 거친 단계에서 이중 소프트맥스를 적용하여 강건한 대응점을 얻는다.
  • SfM에서 투영된 가시 2D-3D 대응을 사용해 거친 매칭 포컬 로스와 미세 2D 좌표 회귀 로스를 결합한 합동 손실로 학습한다。
Figure 1: Comparsion Between Our Method and OnePose [ 48 ] . For low-textured objects that are challenging for OnePose, our method can reconstruct their semi-dense point clouds with more complete geometry and thus achieves more accurate object pose estimation. Green and blue boxes represent ground t
Figure 1: Comparsion Between Our Method and OnePose [ 48 ] . For low-textured objects that are challenging for OnePose, our method can reconstruct their semi-dense point clouds with more complete geometry and thus achieves more accurate object pose estimation. Green and blue boxes represent ground t

실험 결과

연구 질문

  • RQ1키포인트 프리 SfM 파이프라인이 원샷 포즈 추정을 위해 포즈 주석이 달린 제한된 참조 이미지 집합으로부터 정확하고 완전한 반-밀집 3D 객체 모델을 재구성할 수 있는가?
  • RQ2키포인트 프리 재구성 위에 구축된 희소-대-밀도 2D-3D 매칭 네트워크가 표준 데이터셋에서 CAD 모델 비의 기반선과 비교해 경쟁력 있거나 우수한 자세 추정 정확도를 달성하고 CAD 모델 기반 방법에 근접한가?
  • RQ3제안된 방법이 키포인트 기반 방법이 실패하는 저질감 객체에 대해 견고하고, 객체별 훈련 없이 미지의 객체로 일반화할 수 있는가?
  • RQ4제안된 거친-정밀 전략이 실세계 AR 유사 시나리오에서 재구성의 완전성과 자세 추정 정확도에 어떤 영향을 미치는가?

주요 결과

  • OnePose 및 OnePose-LowTexture 데이터셋에서 기존의 원샷 CAD-모델 비의 방법을 큰 차이로 능가합니다.
  • LINEMOD에서 CAD-모델 기반 방법과 유사한 결과를 달성하고, 키포인트 기반 방법이 어려움을 겪는 저질감 객체에서 현저히 더 나은 성능을 보입니다.
  • 제안된 파이프라인은 일부 기반선 대비 현저히 빠르게 실행되며(예: V100에서 512x512 질의의 경우 약 88 ms), 거친-정밀 보정을 통해 더 정확한 반-밀집 재구성을 제공합니다.
  • OnePose-LowTexture에서 메서드는 OnePose 및 HLoc에 비해 상당한 이점을 보여 저질감 영역에 대한 견고함을 강조합니다.
  • LINEMOD에서 다른 원샷 기반선들을 능가하고 CAD 모델 없이 또는 객체별 훈련 없이 인스턴스 수준 방법의 정확도에 가까워집니다.
Figure 2: Overview. 1. For each object, given a reference image sequence $\{\mathbf{I}_{i}\}$ with known object poses $\{\boldsymbol{\xi}_{i}\}$ , our keypoint-free SfM framework reconstructs the semi-dense object point cloud in a coarse-to-fine manner. The coarse reconstruction yields the initial p
Figure 2: Overview. 1. For each object, given a reference image sequence $\{\mathbf{I}_{i}\}$ with known object poses $\{\boldsymbol{\xi}_{i}\}$ , our keypoint-free SfM framework reconstructs the semi-dense object point cloud in a coarse-to-fine manner. The coarse reconstruction yields the initial p

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.