Skip to main content
QUICK REVIEW

[논문 리뷰] Cylinder3D: An Effective 3D Framework for Driving-scene LiDAR Semantic Segmentation

Hui Zhou, Xinge Zhu|arXiv (Cornell University)|2020. 08. 04.
Advanced Neural Network Applications참고 문헌 33인용 수 138
한 줄 요약

Cylinder3D는 원통 기반 분할과 3D 컨볼루션, 그리고 특수한 블록들을 사용하여 LiDAR 포인트 클라우드를 3D에서 직접 처리하면 SemanticKITTI에서 주행 장면 세분화에 대해 최첨단 성능을 보여준다.

ABSTRACT

State-of-the-art methods for large-scale driving-scene LiDAR semantic segmentation often project and process the point clouds in the 2D space. The projection methods includes spherical projection, bird-eye view projection, etc. Although this process makes the point cloud suitable for the 2D CNN-based networks, it inevitably alters and abandons the 3D topology and geometric relations. A straightforward solution to tackle the issue of 3D-to-2D projection is to keep the 3D representation and process the points in the 3D space. In this work, we first perform an in-depth analysis for different representations and backbones in 2D and 3D spaces, and reveal the effectiveness of 3D representations and networks on LiDAR segmentation. Then, we develop a 3D cylinder partition and a 3D cylinder convolution based framework, termed as Cylinder3D, which exploits the 3D topology relations and structures of driving-scene point clouds. Moreover, a dimension-decomposition based context modeling module is introduced to explore the high-rank context information in point clouds in a progressive manner. We evaluate the proposed model on a large-scale driving-scene dataset, i.e. SematicKITTI. Our method achieves state-of-the-art performance and outperforms existing methods by 6% in terms of mIoU.

연구 동기 및 목표

  • 주행-장면 LiDAR 세분화를 위한 2D 투사(projections)의 한계를 평가하고 3D 처리의 이점을 정량화한다.
  • 효과적인 구성을 식별하기 위해 2D 공간과 3D 공간에서의 포인트 표현과 백본(backbones)을 조사한다.
  • 주행-장면 포인트 클라우드에 맞춘 원통 기반 보셀화(voxelization)와 3D CNN 백본을 개발한다.
  • 비대칭 잔차 블록(asymmetric residual block)과 차원 분해 기반 맥락 모델링 모듈을 도입하여 효율성과 맥락 모델링을 개선한다.

제안 방법

  • 원통 좌표계에서 LiDAR 포인트를 보셀화하기 위한 3D 원통 분할을 도입하여 토폴로지를 보존하는 3D 표현을 생성한다.
  • 희소 3D 컨볼루션을 이용한 3D U-Net 백본으로 원통 기반 표현을 처리한다.
  • 표준 잔차 블록을 비대칭 잔차 블록으로 대체하여 직육면체 형태의 주행 객체에 더 잘 맞추고 계산량을 줄인다.
  • 고차원 맥락을 3개의 저차원 구성요소(3x1x1, 1x3x1, 1x1x3)로 분해하고 이를 융합하는 차원 분해 기반 맥락 모델링(DDCM) 모듈을 부착한다.
  • 가중 크로스 엔트로피 손실과 Lovász-softmax 손실을 더해 포인트당 정확도와 mIoU를 최적화하도록 학습한다.
  • 초기 학습률 0.001인 Adam 옵티마이저를 사용한다.

실험 결과

연구 질문

  • RQ1 원통 분할을 통한 3D에서 LiDAR 데이터를 처리하는 것이 주행-장면 세분화를 위한 2D 투사 기반 표현보다 더 우수한가?
  • RQ2 실외 장면의 3D 토폴로지를 포착하는 데 원통 기반 보셀화가 직교 Cartesian 보셀화와 어떻게 비교되는가?
  • RQ3 비대칭 잔차 블록과 차원 분해 맥락 모델링이 전체 성능에 기여하는 바는 무엇인가?
  • RQ4 Cylinder3D가 SemanticKITTI에서 이전 방법들과 비교하여 어떤 성능 이득을 달성하는가?

주요 결과

  • Cylinder3D는 SemanticKITTI에서 최첨단 성능을 달성하고 있으며, mIoU에서 기존 방법들을 상당한 차이로 능가한다.
  • 3D 원통 분할과 3D 컨볼루션은 2D 백본을 투사 기반 표현에 적용한 것에 비해 결과를 크게 개선한다.
  • 표준 잔차 블록을 비대칭 잔차 블록으로 교체하면 약 1.5%의 mIoU 이득을 얻는다.
  • 차원 분해 맥락 모델링 모듈을 도입하면 mIoU가 더 향상되며, 아블레이션 실험에서 상당한 이득이 나타난다.
  • 플립 테스트는 다중 증강 예측을 앙상블할 때 추가로 작은 mIoU 개선을 제공한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.