Skip to main content
QUICK REVIEW

[논문 리뷰] Pyramid R-CNN: Towards Better Performance and Adaptability for 3D Object Detection

Jiageng Mao, Minzhe Niu|arXiv (Cornell University)|2021. 09. 06.
Advanced Neural Network Applications참고 문헌 35인용 수 11
한 줄 요약

Pyramid R-CNN는 포인트 클라우드에서의 희소성과 비균일한 점 분포 문제를 해결하기 위해 새로운 피라미드 RoI 헤드를 도입한 이단계 3D 객체 검출 프레임워크이다. RoI-grid 피라미드, RoI-grid 주의, 밀도 인식 반경 예측을 조합함으로써 희소한 점들로부터의 특징 추출을 향상시켜 KITTI에서 82.08%의 중간 수준 자동차 mAP를 기록하고 Waymo LiDAR 전용 랭킹에서 1위를 달성하여 최신 기술 수준의 성능을 달성하였다.

ABSTRACT

We present a flexible and high-performance framework, named Pyramid R-CNN, for two-stage 3D object detection from point clouds. Current approaches generally rely on the points or voxels of interest for RoI feature extraction on the second stage, but cannot effectively handle the sparsity and non-uniform distribution of those points, and this may result in failures in detecting objects that are far away. To resolve the problems, we propose a novel second-stage module, named pyramid RoI head, to adaptively learn the features from the sparse points of interest. The pyramid RoI head consists of three key components. Firstly, we propose the RoI-grid Pyramid, which mitigates the sparsity problem by extensively collecting points of interest for each RoI in a pyramid manner. Secondly, we propose RoI-grid Attention, a new operation that can encode richer information from sparse points by incorporating conventional attention-based and graph-based point operators into a unified formulation. Thirdly, we propose the Density-Aware Radius Prediction (DARP) module, which can adapt to different point density levels by dynamically adjusting the focusing range of RoIs. Combining the three components, our pyramid RoI head is robust to the sparse and imbalanced circumstances, and can be applied upon various 3D backbones to consistently boost the detection performance. Extensive experiments show that Pyramid R-CNN outperforms the state-of-the-art 3D detection models by a large margin on both the KITTI dataset and the Waymo Open dataset.

연구 동기 및 목표

  • 3D 객체 검출에서 희소하고 비균일하게 분포된 포인트 클라우드의 과제를 해결하기 위해, 특히 먼 거리나 낮은 점 밀도 영역에서의 문제를 해결한다.
  • 두 번째 단계에서 RoI 특징 추출을 향상시켜 이단계 3D 검출기의 강인성과 적응성을 높인다.
  • 다양한 3D 백본(포인트 기반, 볼록체 기반, 포인트-볼록체 기반)과 호환 가능한 일반화 가능한 프레임워크를 설계하여 일관된 성능 향상을 달성한다.
  • RoI 특징 학습 시 풍부한 맥락 정보가 부족한 희소한 포인트 클라우드에서의 성능 저하 문제를 완화한다.
  • 학습 가능한 맥락 인식 특징 추출 메커니즘을 통해 다양한 점 밀도 수준에 동적으로 적응할 수 있도록 한다.

제안 방법

  • 표준 RoI-grid를 다중 수준 피라미드 구조로 확장한 RoI-grid 피라미드를 제안하여 RoI 내외의 중요 포인트를 더 잘 포착함으로써 희소성 문제를 완화한다.
  • 주의 기반 및 그래프 기반 포인트 연산자를 통합한 유일한 공식화인 RoI-grid 주의를 도입하여 RoI 근처의 핵심 포인트를 동적으로 주목한다.
  • 지역적 점 밀도에 따라 각 RoI에 대한 최적의 특징 추출 반경을 예측하는 밀도 인식 반경 예측(DARP) 모듈을 설계한다.
  • RoI-grid 피라미드, RoI-grid 주의, DARP의 세 구성 요소를 하나의 피라미드 RoI 헤드로 통합하여 희소하고 불균형한 조건에서도 특징 표현을 향상시킨다.
  • 다양한 3D 백본(예: PointNet++, VoxelNet, PV-RCNN)에 피라미드 RoI 헤드를 적용하여 일반화 능력과 일관된 성능 향상을 입증한다.
  • 이단계 검출 파이프라인을 적용한다: 첫 번째 단계에서 제안 영역 생성 후, 피라미드 RoI 헤드를 사용해 RoI를 정밀화하여 정렬 및 분류 성능을 향상시킨다.
Figure 1: Statistical results on the KITTI dataset. Blue bars denote the distribution of the number of object points. Orange bars denote the distribution of the number of points gathered by RoIs in Pyramid R-CNN. Our approach can mitigate the sparsity and imbalanced distribution problems of point cl
Figure 1: Statistical results on the KITTI dataset. Blue bars denote the distribution of the number of object points. Orange bars denote the distribution of the number of points gathered by RoIs in Pyramid R-CNN. Our approach can mitigate the sparsity and imbalanced distribution problems of point cl

실험 결과

연구 질문

  • RQ1다중 수준 RoI-grid 구조가 희소한 포인트 클라우드에서의 3D 객체 검출 문제를 효과적으로 완화할 수 있는가?
  • RQ2그래프 기반 및 주의 기반 연산자를 통합한 유일한 주의 메커니즘이 희소한 중요 포인트에서의 특징 추출을 향상시킬 수 있는가?
  • RQ3지역적 점 밀도에 기반해 RoI 반경을 동적으로 조정하면 어려운 낮은 점 수 객체에서 검출 성능을 향상시킬 수 있는가?
  • RQ4제안된 피라미드 RoI 헤드는 다양한 3D 백본 아키텍처(포인트 기반, 볼록체 기반, 하이브리드)에 일반화 가능한가?
  • RQ5KITTI 및 Waymo와 같은 벤치마크 데이터셋에서 희소하고 불균형한 점 분포 조건 하에서 이 프레임워크는 검출 정확도를 어느 정도 향상시키는가?

주요 결과

  • Pyramid R-CNN는 KITTI 데이터셋에서 82.08%의 중간 수준 자동차 mAP를 기록하여 이전 최신 기술 수준의 방법들을 크게 능가하였다.
  • Waymo 오픈 데이터셋에서 Pyramid R-CNN는 LiDAR 전용 방법 중 차량 검출에서 1위를 차지하여 뛰어난 일반화 능력과 강인성을 입증하였다.
  • 제거 분석 결과 각 구성 요소—RoI-grid 피라미드(+1.20% mAP), RoI-grid 주의(+0.51% mAP), DARP(+0.37% mAP)—가 성능 향상에 독립적으로 기여하는 것으로 확인되었다.
  • 피라미드 RoI 헤드는 계산 효율성을 유지하며, KITTI에서 7.86 Hz의 추론 속도를 기록했고, PV-RCNN와 같은 기준 모델보다 약간 느린 수준이었다.
  • 다중 해상도 그리드 배치(Rho_w, Rho_l > 1)를 적용한 RoI-grid 피라미드가 표준 단일 수준 RoI-grid 대비 최대 1.20% 향상된 mAP를 기록하였다.
  • 프레임워크는 백본 간 효과적으로 일반화되었으며, 포인트-볼록체 기반의 Pyramid-PV는 KITTI의 자동차 카테고리에서 88.39%의 AP를 기록하여 PV-RCNN를 1.94% 초월하였다.
Figure 2: The overall architecture. Our Pyramid R-CNN can be plugged on diverse backbones ( e.g . point-based, voxel-based and point-voxel-based networks), which generate 3D proposals and Points of Interest (yellow points) on the stage- $1$ . On the stage- $2$ , we propose the pyramid RoI head that
Figure 2: The overall architecture. Our Pyramid R-CNN can be plugged on diverse backbones ( e.g . point-based, voxel-based and point-voxel-based networks), which generate 3D proposals and Points of Interest (yellow points) on the stage- $1$ . On the stage- $2$ , we propose the pyramid RoI head that

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.