Skip to main content
QUICK REVIEW

[논문 리뷰] Bridging Stereo Geometry and BEV Representation with Reliable Mutual Interaction for Semantic Scene Completion

Bohan Li, Yasheng Sun|arXiv (Cornell University)|2023. 03. 24.
Advanced Vision and Imaging인용 수 6
한 줄 요약

이 논문은 스테레오 기하학과 베이비즈뷰(Bird's-eye-view, BEV) 표현 간을 통합된 점유 기반 아키텍처를 통해 연결하는 새로운 카메라 기반 3D 의미적 장면 완성 프레임워크인 BRGScene을 제안한다. 상호작용 기반 통합(Mutual Interactive Ensemble, MIE) 블록을 도입하여 이중 방향 신뢰성 상호작용 및 이중 볼륨 통합 모듈을 통해 세분화된 신뢰도 인식 특징 융합을 가능하게 하여, SemanticKITTI에서 15.43% mIoU를 달성하며 기존 카메라 전용 방법들보다 유의미한 성능 향상을 이룬다.

ABSTRACT

3D semantic scene completion (SSC) is an ill-posed perception task that requires inferring a dense 3D scene from limited observations. Previous camera-based methods struggle to predict accurate semantic scenes due to inherent geometric ambiguity and incomplete observations. In this paper, we resort to stereo matching technique and bird's-eye-view (BEV) representation learning to address such issues in SSC. Complementary to each other, stereo matching mitigates geometric ambiguity with epipolar constraint while BEV representation enhances the hallucination ability for invisible regions with global semantic context. However, due to the inherent representation gap between stereo geometry and BEV features, it is non-trivial to bridge them for dense prediction task of SSC. Therefore, we further develop a unified occupancy-based framework dubbed BRGScene, which effectively bridges these two representations with dense 3D volumes for reliable semantic scene completion. Specifically, we design a novel Mutual Interactive Ensemble (MIE) block for pixel-level reliable aggregation of stereo geometry and BEV features. Within the MIE block, a Bi-directional Reliable Interaction (BRI) module, enhanced with confidence re-weighting, is employed to encourage fine-grained interaction through mutual guidance. Besides, a Dual Volume Ensemble (DVE) module is introduced to facilitate complementary aggregation through channel-wise recalibration and multi-group voting. Our method outperforms all published camera-based methods on SemanticKITTI for semantic scene completion. Our code is available on https://github.com/Arlo0o/StereoScene.

연구 동기 및 목표

  • 카메라 기반 3D 의미적 장면 완성(semantic scene completion, SSC)에서 내재된 기하학적 모호성과 불완전한 관측 문제를 스테레오 매칭과 BEV 표현을 활용하여 해결하고자 한다.
  • 밀도 높은 3D 장면 완성에서 픽셀 수준의 신뢰성 있는 예측을 위해 스테레오 기하학과 BEV 특징 간의 표현 갭을 해소하고자 한다.
  • 스테레오 및 BEV 특징의 상호 보완적 융합을 통해 가시하지 않은 영역에서의 기하학적 정확도와 의미적 환영 효과를 향상시키고자 한다.
  • 스테레오 볼륨과 BEV 볼륨을 신뢰할 수 있는 상호 특징 상호작용을 통해 통합하는 유일한 프레임워크를 개발하고자 한다.

제안 방법

  • 밀도 높은 3D 스테레오 볼륨과 BEV 특징을 융합하기 위한 통합된 점유 기반 프레임워크인 BRGScene을 제안한다.
  • 스테레오 및 BEV 표현 간 픽셀 수준의 특징 상호작용을 가능하게 하는 상호작용 기반 통합(Mutual Interactive Ensemble, MIE) 블록을 도입한다.
  • 신뢰도 재가중 기능을 통합한 이중 방향 신뢰성 상호작용(Bi-directional Reliable Interaction, BRI) 모듈을 활용하여 특징 정렬을 향상시키고 특징 집합 과정에서 노이즈를 감소시킨다.
  • 채널별 재조정과 다중 그룹 투표를 활용한 이중 볼륨 통합(Dual Volume Ensemble, DVE) 모듈을 도입하여 스테레오 및 BEV 특징의 상호 보완적 융합을 촉진한다.
  • GwcNet을 사용해 밀도 높은 3D 볼륨을 생성하고, BEV 표현 학습을 통해 글로벌 의미적 맥락을 풍부하게 한다.
  • SemanticKITTI 데이터셋에서 교차 엔트로피 손실을 통한 의미 분류와 3D 기하학에 대한 IoU 기반 감독을 사용해 모델을 엔드 투 엔드로 훈련시킨다.
Figure 2: Overall framework of our proposed BRGScene . Given input stereo images, we employ 2D UNet to extract image features. The BEV latent volume and stereo geometric volume are constructed by a BEV Constructor and a Stereo Constructor , respectively. To bridge their representation gap for fine-g
Figure 2: Overall framework of our proposed BRGScene . Given input stereo images, we employ 2D UNet to extract image features. The BEV latent volume and stereo geometric volume are constructed by a BEV Constructor and a Stereo Constructor , respectively. To bridge their representation gap for fine-g

실험 결과

연구 질문

  • RQ13D 의미적 장면 완성에서 스테레오 기하학과 BEV 표현 간을 효과적으로 통합하는 유일한 프레임워크는 어떻게 설계할 수 있는가?
  • RQ2스테레오 및 BEV 특징 간의 상호 신뢰도 인식 특징 상호작용이 장면 완성 정확도에 미치는 영향은 무엇인가?
  • RQ3단일 모odal 입력 대비 이중 볼륨 융합(stereo + BEV)이 기하학적 및 의미적 성능 향상에 얼마나 기여하는가?
  • RQ4제안된 MIE 및 DVE 모듈은 복잡한 실생활 환경에서 단순 연결 또는 단방향 융합보다 우월한 성능을 내는가?
  • RQ5이 프레임워크는 3D 객체 탐지와 같은 SSC 이외의 3D 인식 작업으로도 일반화 가능한가?

주요 결과

  • BRGScene은 SemanticKITTI 검증 세트에서 15.43% mIoU를 달성하여 이전 최고 성능 카메라 기반 방법인 VoxFormer-T 대비 14.5% 상대적 향상을 기록했다.
  • 이중 볼륨 통합(Dual Volume Ensemble, DVE) 모듈은 단순 연결 대비 mIoU를 2.06% 향상시키고 IoU를 3.24% 향상시켜 뚜렷한 기여를 하였다.
  • 이중 방향 신뢰성 상호작용(Bi-directional Reliable Interaction, BRI) 모듈만으로도 기준 모델 대비 mIoU를 2.11% 향상시키고 IoU를 3.36% 향상시켜 상호 지시의 가치를 입증하였다.
  • 스테레오 볼륨을 추가하면 IoU가 3.51 증가하고 mIoU가 1.57 증가하며, BEV 볼륨을 추가하면 IoU가 2.47 증가하고 mIoU가 1.14 증가하여 두 모odal이 상호 보완적 역할을 한다는 점을 확인하였다.
  • 정성적 결과에서는 BRGScene이 교차로나 그림자 영역과 같이 가려진 영역이나 먼 거리의 영역에서 더 정확한 기하학적 구조와 더 나은 환영 효과를 생성하는 것으로 나타났다.
  • nuScenes에서의 초기 BEV 탐지 결과는 BRGScene이 3D 객체 탐지 작업으로도 응용 가능하며, SSC를 넘어서도 넓은 적용 가능성을 지닌다는 것을 시사한다.
Figure 3: The structure of the proposed Bi-directional Reliable Interaction module, which is designed for pixel-level reliable geometry information interaction.
Figure 3: The structure of the proposed Bi-directional Reliable Interaction module, which is designed for pixel-level reliable geometry information interaction.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.