[논문 리뷰] PoinTramba: A Hybrid Transformer-Mamba Framework for Point Cloud Analysis
PoinTramba는 Transformer 기반의 그룹 내 특징 모델링과 Mamba 기반의 그룹 간 처리 를 결합하고, 양방향 중요도 인식 정렬 전략 및 중요도 인식 풀링에 의해 안내되며, 점군 분류 및 세그먼테이션에서 선도적이거나 경쟁력 있는 성과를 달성하고 그룹 간 계산을 선형 시간으로 수행한다.
Point cloud analysis has seen substantial advancements due to deep learning, although previous Transformer-based methods excel at modeling long-range dependencies on this task, their computational demands are substantial. Conversely, the Mamba offers greater efficiency but shows limited potential compared with Transformer-based methods. In this study, we introduce PoinTramba, a pioneering hybrid framework that synergies the analytical power of Transformer with the remarkable computational efficiency of Mamba for enhanced point cloud analysis. Specifically, our approach first segments point clouds into groups, where the Transformer meticulously captures intricate intra-group dependencies and produces group embeddings, whose inter-group relationships will be simultaneously and adeptly captured by efficient Mamba architecture, ensuring comprehensive analysis. Unlike previous Mamba approaches, we introduce a bi-directional importance-aware ordering (BIO) strategy to tackle the challenges of random ordering effects. This innovative strategy intelligently reorders group embeddings based on their calculated importance scores, significantly enhancing Mamba's performance and optimizing the overall analytical process. Our framework achieves a superior balance between computational efficiency and analytical performance by seamlessly integrating these advanced techniques, marking a substantial leap forward in point cloud analysis. Extensive experiments on datasets such as ScanObjectNN, ModelNet40, and ShapeNetPart demonstrate the effectiveness of our approach, establishing a new state-of-the-art analysis benchmark on point cloud recognition. For the first time, this paradigm leverages the combined strengths of both Transformer and Mamba architectures, facilitating a new standard in the field. The code is available at https://github.com/xiaoyao3302/PoinTramba.
연구 동기 및 목표
- Transformer의 정확도와 Mamba의 효율성을 균형 있게 조합하여 점군 분석의 효율성을 촉진한다.
- 그룹 내 의존성과 그룹 간 의존성을 별도로 다루는 하이브리드 아키텍처를 제안한다.
- Mamba 처리에서 발생하는 임의 순서 효과를 완화하기 위해 양방향 중요도 인식 정렬(BIO) 전략을 도입한다.
- 강력한 전역 점군 특징을 형성하기 위한 중요도 인식 풀링 메커니즘을 개발한다.
- 분류 및 세그먼테이션 벤치마크(ScanObjectNN, ModelNet40, ShapeNetPart)에서 효과를 시연한다.
제안 방법
- 점군을 G개의 그룹으로 분할하고 Transformer 인코더를 사용해 그룹 내 임베딩을 추출한다.
- 그룹 임베딩을 그룹 간 Mamba 인코더에 입력해 선형 복잡도로 전역 그룹 간 의존성을 포착한다.
- Mamba 처리 전에 그룹 임베딩의 재정렬을 위해 양방향 중요도 인식 정렬(BIO)을 적용한다.
- 그룹 임베딩에 대한 중요도 점수를 예측해 재정렬을 안내하고 중요도 인식 풀링을 수행해 전역 특징을 형성한다.
- 작업 손실, 중요도 손실, 정합성 손실을 포함한 다중 항 손실로 엔드-투-엔드 학습한다.
실험 결과
연구 질문
- RQ1하이브리드 Transformer-Mamba 아키텍처가 효율성을 유지하면서 분류 및 세그먼테이션에서 단일 아키텍처 모델을 능가할 수 있는가?
- RQ2양방향 중요도 인식 정렬(BIO)이 순서가 무작위인 점군에 대한 Mamba 처리 성능을 개선하는가?
- RQ3중요도 인식 풀링이 재정렬된 그룹 임베딩을 효과적으로 활용해 강건한 전역 특징을 생성하는가?
- RQ4각 구성 요소(Transformer, Mamba, BIO, 정렬, IAP)가 전체 성능에 기여하는 바는 무엇인가?
주요 결과
- PoinTramba는 ScanObjectNN에서 최신 방법과 동등하거나 더 우수한 정확도를 달성합니다(PB-T50-RS 변형: 92.3 ± 0.4; 회전 없이 우리 방법: 92.3 ± 0.2; PB-T50-RS 변형에서 88.9–89.1).
- ModelNet40에서 하이브리드 백본으로 92.7% (±0.1)의 정확도를 달성하여 최근 Transformer/Mamba 방법과 비등하거나 우수합니다.
- ShapeNetPart에서 Inst. mIoU 85.7% (±0.1)로 이전 SOTA 방법과 비슷한 성능을 보입니다.
- 정렬 없이도 그룹 간 Mamba만으로 PointNet++ 대비 정확도가 8.2% 향상되고, 여기에 내부 그룹 Transformer 추가로 0.4% 증가, BIO를 추가하면 2.1% 증가, IAP로 0.5% 증가합니다.
- 정렬과 BIO의 상호작용으로 성능이 더욱 향상되며(정렬: Mamba에 대해 +1.4%, PoinTramba에 대해 +1.7%),
- BIO 정렬은 단방향 또는 무작위 정렬 전략보다 우수하고, IAP는 평균 풀링/최대 풀링보다 더 나은 성능을 보입니다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.