Skip to main content
QUICK REVIEW

[논문 리뷰] RIGA: Rotation-Invariant and Globally-Aware Descriptors for Point Cloud Registration

Hao Yu, Ji Hou|arXiv (Cornell University)|2022. 09. 27.
3D Surveying and Cultural Heritage인용 수 5
한 줄 요약

RIGA는 점군 정합을 위한 새로운 신경 기반 기술을 제안하며, 이는 자체적으로 회전 불변성과 전역 인식 능력을 갖추고 있다. 이는 점 쌍 기능(PPFs)에서 유도된 회전 불변성 있는 局부 기하학적 특징과 비전 트랜스포머를 통한 전역적 구조 및 맥락 정보 통합을 바탕으로 한다. 이는 최신 기술 수준을 초월하는 성능을 달성하며, 대량의 회전 조건 하에서 ModelNet40에서 상대적 회전 오차를 8° 감소시키고, 3DLoMatch에서 특징 매칭 재현율을 최소 5%p 향상시킨다.

ABSTRACT

Successful point cloud registration relies on accurate correspondences established upon powerful descriptors. However, existing neural descriptors either leverage a rotation-variant backbone whose performance declines under large rotations, or encode local geometry that is less distinctive. To address this issue, we introduce RIGA to learn descriptors that are Rotation-Invariant by design and Globally-Aware. From the Point Pair Features (PPFs) of sparse local regions, rotation-invariant local geometry is encoded into geometric descriptors. Global awareness of 3D structures and geometric context is subsequently incorporated, both in a rotation-invariant fashion. More specifically, 3D structures of the whole frame are first represented by our global PPF signatures, from which structural descriptors are learned to help geometric descriptors sense the 3D world beyond local regions. Geometric context from the whole scene is then globally aggregated into descriptors. Finally, the description of sparse regions is interpolated to dense point descriptors, from which correspondences are extracted for registration. To validate our approach, we conduct extensive experiments on both object- and scene-level data. With large rotations, RIGA surpasses the state-of-the-art methods by a margin of 8\degree in terms of the Relative Rotation Error on ModelNet40 and improves the Feature Matching Recall by at least 5 percentage points on 3DLoMatch.

연구 동기 및 목표

  • 기존 신경 기반 기술이 회전 불변성 없는 기반 구조로 인해 큰 회전 조건에서 성능이 저하되는 문제를 해결한다.
  • 국소 기하학적 특징을 넘어서 3D 전역적 구조적 맥락을 통합함으로써 기술의 특징 구별 능력을 향상시킨다.
  • 희박한 샘플링 조건에서의 반복성 문제를 해결하기 위해 강력한 기술을 활용해 코arse-to-fine 대응 매칭을 가능하게 한다.
  • 임의의 회전에 대해 불변성을 유지하면서도 높은 구별 능력을 유지하는 기술을 설계한다.
  • 어려운 회전 조건 하에서도 물체 수준 및 장면 수준의 벤치마크에서 뛰어난 정합 성능을 달성한다.

제안 방법

  • 점군에서 희박한 국소 영역을 추출하고, 국소 기하학적 특징을 표현하기 위해 점 쌍 기능(PPFs)을 계산한다.
  • PPFs를 회전 불변성 있는 국소 기하 기술로 변환하기 위해 회전 불변성 인코딩 모듈을 적용한다.
  • 전체 3D 구조를 표현하기 위해 전역 PPF 서명을 구성하여 전역적 구조 인식 능력을 제공한다.
  • 비전 트랜스포머(ViT)를 사용해 전역 PPF 서명에서 구조적 기술을 학습함으로써 국소 기술에 전역 맥락을 통합한다.
  • ViT 인코더의 자기주의 기반 메커니즘을 통해 장면 수준의 기하학적 맥락을 국소 기술에 전역적으로 통합한다.
  • 코arse-to-fine 정합 파이프라인에서 효율적이고 신뢰할 수 있는 대응 매칭을 위해 희박한 기술을 밀도 있는 점 기반 기술로 보간한다.
Figure 1: Feature Matching Recall (FMR) on 3DLoMatch [ 2 ] (x-axis) and Rotated 3DLoMatch (y-axis). Methods that only encode local geometry are marked as blue, while approaches with global awareness are drawn in red. The performance drop from the original (x-axis) to the rotated (y-axis) benchmark f
Figure 1: Feature Matching Recall (FMR) on 3DLoMatch [ 2 ] (x-axis) and Rotated 3DLoMatch (y-axis). Methods that only encode local geometry are marked as blue, while approaches with global awareness are drawn in red. The performance drop from the original (x-axis) to the rotated (y-axis) benchmark f

실험 결과

연구 질문

  • RQ1고도로 구별 가능한 성능을 유지하면서도 본질적으로 회전 불변성을 갖춘 신경 기술을 설계할 수 있는가?
  • RQ2전역적 구조적 및 맥락적 정보를 통합할 경우, 큰 회전 조건 하에서도 기술의 내성적 성능 향상은 어떻게 이루어지는가?
  • RQ3순수하게 국소 기반 기술에 비해 전역 인식 능력이 있는 기술이 얼마나 성능 저하를 줄이는가?
  • RQ4회전 불변성과 전역 인식 능력을 갖춘 기술이 물체 수준 및 장면 수준의 정합 작업에서 최신 기술 수준을 초월할 수 있는가?
  • RQ5실제 LiDAR 데이터에서 흔히 발생하는 정규 벡터 추정의 열악한 조건 하에서도 제안된 기술은 얼마나 내성적인가?

주요 결과

  • 대량의 회전 조건 하에서 ModelNet40에서 RIGA는 최신 기술 수준의 방법 대비 상대적 회전 오차를 8° 감소시켰다.
  • 3DLoMatch에서 RIGA는 조건이 변형된 테스트 환경에서도 특징 매칭 재현율을 최소 5%p 향상시켰다.
  • KITTI에서 RIGA는 정합 재현율 99.1%를 달성하여, 더 높은 계산 비용을 지닌 ViT 기반 아키텍처임에도 불구하고 대부분의 베이스라인을 능가했다.
  • 회전된 3DLoMatch에서 RIGA는 성능 저하가 0.6%에 그쳐 비교된 모든 방법 중에서 가장 낮은 수준을 보이며 뛰어난 회전 내성성의 우수성을 입증했다.
  • 실외 KITTI 스캔에서 정규 벡터 추정이 열악한 조건 하에서도 RIGA는 모든 메트릭에서 최신 기술 수준의 방법과 유사한 성능을 유지했다.
  • 절단 분석 결과, 전역적 구조 인코딩과 전역 맥락 통합 모두 성능 향상에 기여하며, ViT 기반의 주의 메커니즘이 맥락 통합의 핵심 역할을 한다는 것이 확인되었다.
Figure 2: Illustration of the Inherent Rotational Invariance and Distinctiveness of RIGA. In (a), an arbitrary rotation is applied to the input scan. 1) Rotational Invariance : In (b), (c) and (d), local, global and point descriptors from untrained RIGA are visualized by t-SNE [ 5 ] , respectively.
Figure 2: Illustration of the Inherent Rotational Invariance and Distinctiveness of RIGA. In (a), an arbitrary rotation is applied to the input scan. 1) Rotational Invariance : In (b), (c) and (d), local, global and point descriptors from untrained RIGA are visualized by t-SNE [ 5 ] , respectively.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.