Skip to main content
QUICK REVIEW

[논문 리뷰] FusionFormer: A Multi-sensory Fusion in Bird's-Eye-View and Temporal Consistent Transformer for 3D Object Detection

Chunyong Hu, Hang Zheng|arXiv (Cornell University)|2023. 09. 11.
Advanced Neural Network Applications인용 수 5
한 줄 요약

FusionFormer는 2D 이미지 및 3D 볼록체 특징을 직접 융합하기 위해 균일한 샘플링 전략을 사용하여 명시적인 Bird's-Eye-View (BEV) 변환을 피하는 엔드 투 엔드 다중 감각 융합 트랜스포머를 제안한다. 이는 탈형 변형 주의 메커니즘과 잔차 구조를 활용하여 강건성을 확보한다. FusionFormer는 테스트 시각 증강 없이 nuScenes에서 72.6% mAP 및 75.1% NDS를 기록하여 최신 기술 수준을 달성한다.

ABSTRACT

Multi-sensor modal fusion has demonstrated strong advantages in 3D object detection tasks. However, existing methods that fuse multi-modal features require transforming features into the bird's eye view space and may lose certain information on Z-axis, thus leading to inferior performance. To this end, we propose a novel end-to-end multi-modal fusion transformer-based framework, dubbed FusionFormer, that incorporates deformable attention and residual structures within the fusion encoding module. Specifically, by developing a uniform sampling strategy, our method can easily sample from 2D image and 3D voxel features spontaneously, thus exploiting flexible adaptability and avoiding explicit transformation to the bird's eye view space during the feature concatenation process. We further implement a residual structure in our feature encoder to ensure the model's robustness in case of missing an input modality. Through extensive experiments on a popular autonomous driving benchmark dataset, nuScenes, our method achieves state-of-the-art single model performance of 72.6% mAP and 75.1% NDS in the 3D object detection task without test time augmentation.

연구 동기 및 목표

  • 기존의 다중 감각 융합 방법에서 3D 볼록체 특징을 BEV 공간으로 명시적으로 변환함으로써 발생하는 성능 저하 문제를 해결하기 위해.
  • 단일 깊이 추정 또는 Z축 압축에 의존하지 않고 2D 이미지 및 3D 포인트 클라우드 특징을 민첩하게 직접 융합할 수 있도록 하기 위해.
  • 이미지 특징 변환을 위한 깊이 참조로 희소 포인트 클라우드 특징을 사용하여 특징 표현을 향상시키기 위해.
  • 모달리티 입력이 누락된 상황에서도 잔차 학습을 통해 모델의 강건성을 향상시키기 위해.
  • 이전 프레임의 통합을 위해 플러그 앤 플레이 방식의 시간적 융합 모듈을 통합하여 시간적 일관성을 확보하기 위해.

제안 방법

  • 2D 이미지 특징과 3D 볼록체 특징을 동시에 샘플링할 수 있는 균일한 샘플링 전략을 도입하여 명시적인 BEV 변환을 피한다.
  • 탈형 변형 주의 메커니즘을 활용하여 BEV 쿼리가 3D 볼록체에서 유도된 깊이 參조를 기반으로 이미지 및 포인트 클라우드 특징과 상호작용하도록 한다.
  • 모달리티가 누락된 추론 상황에서도 성능을 유지하기 위해 융합 인코더에 잔차 구조를 통합한다.
  • 이전 프레임의 BEV 특징을 집계하여 시간적 일관성을 향상시키는 시간적 융합 모듈을 설계한다.
  • 객체 쿼리 정밀화 및 검출 헤드 예측을 위해 다중 헤드 주의 기반 트랜스포머 디코더를 활용한다.
  • LiDAR 특징을 단일 깊이 추정 출력으로 대체함으로써 카메라 전용 3D 검출을 지원하여 방법의 유연성을 입증한다.
Figure 1: Comparison between state-of-the-art methods and our FusionFormer. (a) In BEVFusion-based methods, the camera features and points features are transformed into BEV space and fused with concatenation. (b) In CMT, the points voxel features are first compressed into BEV features, then are enco
Figure 1: Comparison between state-of-the-art methods and our FusionFormer. (a) In BEVFusion-based methods, the camera features and points features are transformed into BEV space and fused with concatenation. (b) In CMT, the points voxel features are first compressed into BEV features, then are enco

실험 결과

연구 질문

  • RQ12D 이미지 및 3D 볼록체 특징을 그들의 원천 형태 그대로 직접 융합하는 것이 BEV 변환 기반 융합보다 3D 객체 검출에서 더 우수한 성능을 내는가?
  • RQ23D 볼록체 특징을 깊이 참조로 사용할 경우 BEV 공간에서의 이미지 특징 변환 정확도가 향상되는가?
  • RQ3통합 샘플링 전략이 2D 및 3D와 같은 이질적인 모달리티를 단일 융합 모듈에서 효과적으로 처리할 수 있는가?
  • RQ4융합 인코더에서의 잔차 학습이 모달리티 입력 누락에 대한 강건성에 어떤 영향을 미치는가?
  • RQ5이전 BEV 특징의 시간적 융합이 검출 일관성과 성능 향상에 얼마나 기여하는가?

주요 결과

  • FusionFormer는 테스트 시각 증강 없이 단일 모델로 nuScenes 벤치마크에서 72.6% mAP 및 75.1% NDS를 기록하여 최신 기술 수준의 성능을 달성한다.
  • LiDAR 입력으로 BEV 특징 대신 볼록체 특징을 사용할 경우 mAP가 1.4%p 향상되고 mATE가 1.3p 감소하여 3D 구조적 정보의 보존이 향상됨을 시사한다.
  • 제안된 융합 모듈은 단순한 덧셈 및 연결 방식보다 mAP 3.5p, NDS 2.8p 높은 성능을 기록하여 뛰어난 특징 융합 능력을 입증한다.
  • 모달리티 누락 상황에서도 뛰어난 성능 유지를 보이며, 이미지 데이터가 누락되었을 경우 65.2% mAP 및 69.8% NDS, LiDAR 데이터가 누락되었을 경우 64.1% mAP 및 68.5% NDS를 기록하여 강건성을 확인한다.
  • 시각화 결과 FusionFormer는 BEVFormer 대비 목표 물체, 특히 먼 곳이나 희소한 물체에 대해 더 집중적인 BEV 특징을 생성하는 것으로 나타났다.
  • 단일 카메라 입력만을 사용하는 FusionFormer의 카메라 전용 버전은 단일 깊이 추정을 활용하여 68.9% mAP 및 71.3% NDS를 기록하여 카메라 입력만으로도 뛰어난 성능을 내는 것으로 확인되었다.
Figure 2: (a) The framework of the FusionFormer. The LiDAR point cloud and multi-view images are processed separately in their respective backbone networks to extract voxel features and image features. These features are then inputted into a multi-modal fusion encoder (MMFE) to generate the fused BE
Figure 2: (a) The framework of the FusionFormer. The LiDAR point cloud and multi-view images are processed separately in their respective backbone networks to extract voxel features and image features. These features are then inputted into a multi-modal fusion encoder (MMFE) to generate the fused BE

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.