[논문 리뷰] Fusion-Mamba for Cross-modality Object Detection
이 논문은 SSCS와 DSSF를 사용하여 숨겨진 상태 공간에서 RGB와 IR 특징을 융합하는 Fusion-Mamba 블록을 제안하고, 다중 데이터셋에서 교차 모달 물체 탐지에서 최첨단 성능을 달성합니다.
Cross-modality fusing complementary information from different modalities effectively improves object detection performance, making it more useful and robust for a wider range of applications. Existing fusion strategies combine different types of images or merge different backbone features through elaborated neural network modules. However, these methods neglect that modality disparities affect cross-modality fusion performance, as different modalities with different camera focal lengths, placements, and angles are hardly fused. In this paper, we investigate cross-modality fusion by associating cross-modal features in a hidden state space based on an improved Mamba with a gating mechanism. We design a Fusion-Mamba block (FMB) to map cross-modal features into a hidden state space for interaction, thereby reducing disparities between cross-modal features and enhancing the representation consistency of fused features. FMB contains two modules: the State Space Channel Swapping (SSCS) module facilitates shallow feature fusion, and the Dual State Space Fusion (DSSF) enables deep fusion in a hidden state space. Through extensive experiments on public datasets, our proposed approach outperforms the state-of-the-art methods on $m$AP with 5.9% on $M^3FD$ and 4.9% on FLIR-Aligned datasets, demonstrating superior object detection performance. To the best of our knowledge, this is the first work to explore the potential of Mamba for cross-modal fusion and establish a new baseline for cross-modality object detection.
연구 동기 및 목표
- RGB와 IR 데이터 간의 모달리티 차이를 해소함으로써 강인한 교차 모달 물체 탐지를 촉진한다.
- 모달리티 간 특징 융합을 향상시키기 위해 숨겨진 상태 상호작용 메커니즘을 개발한다.
- 과도한 계산 없이 표현 일관성을 개선하는 효율적인 융합 모듈을 설계한다.
- 여러 백본에 걸친 공개 RGB-IR 데이터셋에서 최첨단 성능을 입증한다.
제안 방법
- 두 모듈로 구성된 Fusion-Mamba blocks (FMB): 얕은 융합을 위한 State Space Channel Swapping (SSCS) 및 깊은 융합을 위한 Dual State Space Fusion (DSSF)을 숨겨진 상태 공간에서.
- 교차 모달 특징을 숨겨진 상태 공간으로 매핑하여 상호 작용을 가능하게 하고 모달리티 차이를 줄인다.
- SSCS는 채널 스와핑을 수행하고 얕은 교차 모달 상호 작용을 위해 Visual State Space (VSS) 블록을 적용한다.
- DSSF는 깊은 교차 모달 융합을 가능하게 하는 게이팅 메커니즘으로 특징을 숨겨진 상태 공간으로 투영하고, 그 후 다시 투영하여 잔차 연결을 수행한다.
- 향상된 RGB 및 IR 특징을 더하기 방식으로 융합하여 탐지 헤드의 YOLOv5/YOLOv8 백본에서 사용되는 융합 특징을 형성한다.
실험 결과
연구 질문
- RQ1RGB와 IR 특징 간의 교차 모달 차이가 융합 중 어떻게 완화되어 물체 탐지 정확도를 향상시킬 수 있는가?
- RQ2숨겨진 상태 공간 융합 접근 방식이 RGB-IR 데이터에 대해 공간 중심 또는 트랜스포머 기반 융합을 능가할 수 있는가?
- RQ3Fusion-Mamba를 트랜스포머 기반 융합 방법과 비교했을 때 정확도와 계산 비용의 트레이드오프는 무엇인가?
주요 결과
- RGB-IR 데이터로 Fusion-Mamba는 이전 방법 대비 M3FD에서 mAP가 5.9% 포인트, FLIR-Aligned에서 4.9% 포인트의 최첨단 개선을 달성했다.
- LLVIP에서 YOLOv5 백본을 이용한 Fusion-Mamba는 62.8 mAP 및 96.8 mAP50을 달성했고, YOLOv8 백본으로는 64.3 mAP 및 97.0 mAP50에 도달했다.
- YOLOv8을 사용한 Fusion-Mamba는 LLVIP, M3FD, FLIR-Aligned 데이터셋 전반에 걸쳐 백본 간에 경쟁적 기초 모델들보다 일관되게 더 나은 성능을 보인다.
- 제거 연구에서 SSCS 또는 DSSF를 제거하면 mAP가 현저한 감소를 보였으며, 이는 효과적인 교차 모달 융합을 위해 두 모듈의 필요성을 보여준다.
- 트랜스포머 기반 융합과 비교하여 Fusion-Mamba는 시간 복잡도가 낮고 가시적인 속도 향상을 제공한다(예: 동일 설정에서 추론이 7~19 ms 더 빠름).
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.