Skip to main content
QUICK REVIEW

[논문 리뷰] RoboBEV: Towards Robust Bird's Eye View Perception under Corruptions

Shaoyuan Xie, Lingdong Kong|arXiv (Cornell University)|2023. 04. 13.
Video Surveillance and Tracking Methods인용 수 4
한 줄 요약

이 논문은 환경적( dense fogs, snow), 센서 유도적( motion blur, color quantization) 및 시간적 교란( camera crash, frame loss)을 포함한 8종의 자연적 손상 상황에서 카메라 기반 뷰(鳥瞰図, BEV) 인식 모델의 강건성 평가를 위한 종합적 벤치마크인 RoboBEV를 소개한다. 연구 결과, 사전 훈련 및 깊이 없는 BEV 변환 기법이 분포 외 강건성 향상에 기여하며, 장거리 시간적 모델링이 손상 상황에서 성능 향상에 크게 기여하는 것으로 나타났다.

ABSTRACT

The recent advances in camera-based bird's eye view (BEV) representation exhibit great potential for in-vehicle 3D perception. Despite the substantial progress achieved on standard benchmarks, the robustness of BEV algorithms has not been thoroughly examined, which is critical for safe operations. To bridge this gap, we introduce RoboBEV, a comprehensive benchmark suite that encompasses eight distinct corruptions, including Bright, Dark, Fog, Snow, Motion Blur, Color Quant, Camera Crash, and Frame Lost. Based on it, we undertake extensive evaluations across a wide range of BEV-based models to understand their resilience and reliability. Our findings indicate a strong correlation between absolute performance on in-distribution and out-of-distribution datasets. Nonetheless, there are considerable variations in relative performance across different approaches. Our experiments further demonstrate that pre-training and depth-free BEV transformation has the potential to enhance out-of-distribution robustness. Additionally, utilizing long and rich temporal information largely helps with robustness. Our findings provide valuable insights for designing future BEV models that can achieve both accuracy and robustness in real-world deployments.

연구 동기 및 목표

  • 실세계의 손상 상황에서 카메라 기반 BEV 인식 모델의 강건성 이해에 있어 중요한 격차를 보완하기 위해.
  • 자율주행 시나리오에서 흔히 발생하는 다양한 자연적 손상 상황에서 기존 BEV 모델의 성능을 평가하기 위해.
  • 분포 내 성능에만 의존하지 않고도 강건성을 향상시키는 아키텍처 및 훈련 전략을 규명하기 위해.
  • 안전 중심의 배치에서 정확하고 신뢰할 수 있는 미래의 BEV 모델 설계를 위한 실질적 통찰를 제공하기 위해.

제안 방법

  • Bright, Dark, Fog, Snow, Motion Blur, Color Quant, Camera Crash, Frame Lost의 8종의 고유한 손상 유형을 포함한 새로운 벤치마크 슈트인 RoboBEV를 제안한다.
  • nuScenes를 포함한 표준 BEV 검출 벤치마크에 대해 세 가지 심각도 수준의 손상을 적용하여 실제 분포 이탈 상황을 시뮬레이션한다.
  • 청결 상태와 손상 상태 양쪽에서 26종의 최신 카메라 기반 BEV 인식 모델을 평가하여 강건성 수준을 분석한다.
  • 포prehensive 평가를 위해 NuScenes Detection Score(NDS), mAP, 그리고 정위치 오차(mATE, mASE 등)와 같은 표준화된 지표를 활용한다.
  • 사전 훈련, 깊이 없는 BEV 변환, 시간적 모델링의 영향을 분석하기 위해 다양한 손상 유형에 대해 모델 변종 간 비교 분석을 수행한다.
  • 성능 저하 패턴과 모델 간 강건성 패턴을 시각화하기 위해 레이더 차트와 통계 분석을 활용한다.
Figure 1: The radar charts of existing BEV detectors’ nuScenes Detection Score (NDS) [ 3 ] under eight corruption types. We observe diverse behaviors of different models even with competitive “clean” performance. The NDS is normalized across all the benchmarking BEV models to lie between 0.1 and 1.
Figure 1: The radar charts of existing BEV detectors’ nuScenes Detection Score (NDS) [ 3 ] under eight corruption types. We observe diverse behaviors of different models even with competitive “clean” performance. The NDS is normalized across all the benchmarking BEV models to lie between 0.1 and 1.

실험 결과

연구 질문

  • RQ1기존 BEV 모델은 다양한 자연적 손상 상황에서 어떻게 성능을 보이며, 청결 상태 성능과 손상 상황에서의 강건성 간 강한 상관관계가 존재하는가?
  • RQ2사전 훈련이나 깊이 없는 BEV 변환과 같은 아키텍처나 훈련 구성 요소 중에서 분포 외 강건성 향상에 가장 크게 기여하는 것은 무엇인가?
  • RQ3장기적이고 풍부한 시간적 정보를 활용할 경우, 프레임 손실이나 카메라 고장과 같은 시간적 손상에 대한 모델의 내성 강화 정도는 어느 정도인가?
  • RQ4특정 손상 유형이 모델 성능을 비례적으로 크게 떨어뜨리며, 이는 모델 아키텍처에 따라 다를 수 있는가?

주요 결과

  • 분포 내 성능과 손상 상황에서의 강건성 간 강한 상관관계가 존재하지만, 상대적 강건성은 절대 성능 순위와 항상 일치하지는 않는다.
  • 사전 훈련 및 깊이 없는 BEV 변환을 사용하는 모델은 모든 손상 유형에서 뚜렷한 강건성 향상을 보이며, 이는 핵심적인 설계 이점임을 시사한다.
  • 더 길고 풍부한 시간적 모델링은 특히 Frame Lost 및 Camera Crash와 같은 시간적 손상 상황에서 강건성 향상에 크게 기여한다.
  • Frame Lost 및 Camera Crash와 같은 시간적 손상은 가장 심각한 성능 저하를 유발하며, 일부 모델은 NDS가 50% 이상 감소하는 경우가 있다.
  • Snow 및 Fog와 같은 환경적 손상은 상당한 성능 저하를 유발하며, 극한의 경우(예: SOLOFusion의 Snow 상황) NDS가 0.11까지 하락하는 경우가 있다.
  • Color Quantization 및 Motion Blur 역시 mAP 및 정위치 오차 지표에서 뚜렷한 성능 저하를 유발하며, 이는 시각적 정밀도에 대한 민감성을 드러낸다.
Figure 2: Corruption examples from the RoboBEV benchmark. Left: Corruption taxonomy. Right: Temporal corruptions. Camera Crash drop fixed set of images along timestamps; Frame Lost randomly drop frames along timestamps. More examples are in Appendix 9.3 .
Figure 2: Corruption examples from the RoboBEV benchmark. Left: Corruption taxonomy. Right: Temporal corruptions. Camera Crash drop fixed set of images along timestamps; Frame Lost randomly drop frames along timestamps. More examples are in Appendix 9.3 .

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.