[논문 리뷰] FlashOcc: Fast and Memory-Efficient Occupancy Prediction via Channel-to-Height Plugin
FlashOcc는 Bird's Eye View (BEV) 특징 공간에서 3D 컨볼루션을 2D 컨볼루션으로 대체하고, 채널에서 높이로의 변환을 적용하여 3D 오염도 로짓을 생성함으로써 빠르고 메모리 효율적인 3D 오염도 예측을 위한 플러그 앤 플레이 프레임워크를 제안한다. nuScenes에서 기준 모델 대비 58.7% 빠른 추론 속도, 68.8% 낮은 메모리 사용량, 최대 6.1 mIoU 향상을 달성하여 최신 기술 수준을 확립한다.
Given the capability of mitigating the long-tail deficiencies and intricate-shaped absence prevalent in 3D object detection, occupancy prediction has become a pivotal component in autonomous driving systems. However, the procession of three-dimensional voxel-level representations inevitably introduces large overhead in both memory and computation, obstructing the deployment of to-date occupancy prediction approaches. In contrast to the trend of making the model larger and more complicated, we argue that a desirable framework should be deployment-friendly to diverse chips while maintaining high precision. To this end, we propose a plug-and-play paradigm, namely FlashOCC, to consolidate rapid and memory-efficient occupancy prediction while maintaining high precision. Particularly, our FlashOCC makes two improvements based on the contemporary voxel-level occupancy prediction approaches. Firstly, the features are kept in the BEV, enabling the employment of efficient 2D convolutional layers for feature extraction. Secondly, a channel-to-height transformation is introduced to lift the output logits from the BEV into the 3D space. We apply the FlashOCC to diverse occupancy prediction baselines on the challenging Occ3D-nuScenes benchmarks and conduct extensive experiments to validate the effectiveness. The results substantiate the superiority of our plug-and-play paradigm over previous state-of-the-art methods in terms of precision, runtime efficiency, and memory costs, demonstrating its potential for deployment. The code will be made available.
연구 동기 및 목표
- 자율주행 시스템에서 3D 볼륨 기반 오염도 예측의 높은 메모리 및 계산 비용을 해결한다.
- 칩 내부 배치를 저해하는 3D 컨볼루션 및 트랜스포머 기반 모듈의 한계를 극복한다.
- 다양한 기존 오염도 예측 모델과 호환되는 일반적인 목적의 플러그 앤 플레이 솔루션을 개발한다.
- 추론 지연과 메모리 소비를 극적으로 줄이면서도 높은 정확도를 유지한다.
- 아키텍처 재설계 없이 다양한 하드웨어 플랫폼에 효율적으로 배포 가능하게 한다.
제안 방법
- 기존 볼륨 기반 오염도 모델의 3D 컨볼루션 레이어를 Bird's Eye View (BEV) 특징에 작동하는 효율적인 2D 컨볼루션 레이어로 대체한다.
- 각 BEV 픽셀이 피라미드 수준의 높이 정보를 인코딩하는 BEV 표현에서 공간적 및 채널 수준의 정보를 유지한다.
- 플랫티드된 BEV 특징을 3D 볼륨 수준의 오염도 로짓으로 재구성하기 위해 채널에서 높이로의 변환을 도입한다.
- 서브픽셀 컨볼루션 원리(채널 재배열)를 활용하여 3D 공간으로의 효율적 업샘플링을 가능하게 하며, 3D 연산을 수행하지 않는다.
- 기존 모델에 플러그인으로 적용하여 BEV 인코더와 오염도 헤드 컴ponent만 수정한다.
- 프레임 간 특징 일관성을 유지함으로써 시간적 융합 모듈과의 호환성을 확보한다.

실험 결과
연구 질문
- RQ1정확도를 유지하거나 향상시키면서 플러그 앤 플레이 모듈이 3D 컨볼루션을 오염도 예측에서 대체할 수 있는가?
- RQ2채널에서 높이로의 변환은 2D BEV 특징에서 3D 공간 정보를 어느 정도 유지할 수 있는가?
- RQ33D 컨볼루션 기반 기준 모델 대비 제안된 방법이 메모리 소비와 추론 지연을 얼마나 줄이는가?
- RQ4이 방법은 다양한 백본 아키텍처와 시간 모델링 전략에 대해 일반화 가능한가?
- RQ5자원 제약이 있는 엣지 디바이스에서 배포 가능한 상태로 높은 성능을 달성할 수 있는가?
주요 결과
- BEVDetOcc에서 BEV 인코더와 오염도 헤드의 추론 시간이 58.7% 감소하여 7.5 ms에서 3.1 ms로 줄어들었다.
- BEVDetOcc에서 추론 중 메모리 소비가 68.8% 감소하여 398 MiB에서 124 MiB로 감소했다.
- nuScenes 벤치마크에서 시간적 융합을 사용할 경우 기준 모델 UniOcc 대비 6.1 mIoU 향상을 달성했다.
- 비시간적 및 시간적 BEVDetOcc 변종에서 각각 원래 방법 대비 mIoU가 0.8점과 1.7점 향상되었다.
- 비시간적 설정에서는 학습 기간이 50% 감소(64에서 32 에포크), 시간적 설정에서는 42% 감소(144에서 84 에포크)되었다.
- 시각화를 통한 확인과 일관된 성능 향상으로 채널에서 높이로의 변환이 높이 정보를 성공적으로 유지함을 입증했다.

더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.