Skip to main content
QUICK REVIEW

[논문 리뷰] BRAU-Net++: U-Shaped Hybrid CNN-Transformer Network for Medical Image Segmentation

Libin Lan, Pengzhou Cai|arXiv (Cornell University)|2024. 01. 01.
Radiomics and Machine Learning in Medical Imaging인용 수 24
한 줄 요약

BRAU-Net++은 코어 블록으로 Bi-Level Routing Attention(BRA)과 채널-공간 건너뛰기 연결(SCCSA)을 사용하는 U자형 CNN-Transformer 아키텍처로, 다수의 데이터셋에서 효율적이고 정확한 의학 영상 분할을 달성합니다.

ABSTRACT

Accurate medical image segmentation is essential for clinical quantification, disease diagnosis, treatment planning and many other applications. Both convolution-based and transformer-based u-shaped architectures have made significant success in various medical image segmentation tasks. The former can efficiently learn local information of images while requiring much more image-specific inductive biases inherent to convolution operation. The latter can effectively capture long-range dependency at different feature scales using self-attention, whereas it typically encounters the challenges of quadratic compute and memory requirements with sequence length increasing. To address this problem, through integrating the merits of these two paradigms in a well-designed u-shaped architecture, we propose a hybrid yet effective CNN-Transformer network, named BRAU-Net++, for an accurate medical image segmentation task. Specifically, BRAU-Net++ uses bi-level routing attention as the core building block to design our u-shaped encoder-decoder structure, in which both encoder and decoder are hierarchically constructed, so as to learn global semantic information while reducing computational complexity. Furthermore, this network restructures skip connection by incorporating channel-spatial attention which adopts convolution operations, aiming to minimize local spatial information loss and amplify global dimension-interaction of multi-scale features. Extensive experiments on three public benchmark datasets demonstrate that our proposed approach surpasses other state-of-the-art methods including its baseline: BRAU-Net under almost all evaluation metrics. We achieve the average Dice-Similarity Coefficient (DSC) of 82.47, 90.10, and 92.94 on Synapse multi-organ segmentation, ISIC-2018 Challenge, and CVC-ClinicDB, as well as the mIoU of 84.01 and 88.17 on ISIC-2018 Challenge and CVC-ClinicDB, respectively.

연구 동기 및 목표

  • CNN 로컬 모델링과 Transformer 기반 글로벌 의존성의 결합으로 정확한 의학 영상 분할을 촉진합니다.
  • Bi-Level Routing Attention(BRA)을 통해 장거리 어텐션의 계산 복잡도를 감소시킵니다.
  • SCCSA를 통한 채널-공간 어텐션 주도 스킵 연결로 공간 세부 정보를 보존합니다.
  • SOTA 방법과 비교하여 다양한 의학 영상 데이터셋에서 성능 향상을 보여줍니다.

제안 방법

  • Core로 BiFormer 블록을 사용한 U자형 인코더-디코더를 구축하고, BRA를 지역 간(attention)과 토큰 간(attention)으로 사용합니다.
  • 패치 임베딩/병합 및 패치 확장을 통해 대칭 아키텍처를 형성하는 계층적 인코더-디코더를 적용합니다.
  • SCCSA를 사용해 채널과 공간 주의력을 컨볼루션으로 구현하여 차원 간 상호작용을 강화하도록 스킵 연결을 재설계합니다.
  • 깊이 방향 컨볼루션을 사용해 위치 정보를 인코딩하고, 데이터셋 특성 가중치를 갖는 Dice + Cross-Entropy 하이브리드 손실을 적용합니다.
  • 전처리 및 패치 기반 토큰화를 이용해 O((HW)^(4/3)) 복잡도의 효율적이고 쿼리 인식형 희소 어텐션을 가능하게 합니다.
  • Synapse, ISIC-2018, CVC-ClinicDB에서 데이터 증강과 표준 최적화 스케줄로 학습합니다.

실험 결과

연구 질문

  • RQ1BRAU-Net++이 다양한 모달리티에서 최첨단 CNN-, Transformer-, 하이브리드 기반 의학 영상 분할 방법을 능가할 수 있는가?
  • RQ2Bi-Level Routing Attention이 계산 및 메모리 감소로 효과적인 장거리 의존성 모델링을 가능하게 하는가?
  • RQ3SCCSA 스킵 연결 전략이 디코딩 중 공간 세부 정보 보존과 다중 스케일 특성 통합을 개선하는가?

주요 결과

  • BRAU-Net++는 ISIC-2018에서 DSC 82.47%, mIoU 84.01% 등의 평균 지표를 포함한 벤치마크 전반에서 높은 분할 성능을 달성합니다(또한 CVC-ClinicDB에서도 유사한 수치를 보입니다).
  • TransUNet 및 Swin-Unet과 비교해 BRAU-Net++이 DSC를 각각 4.49% 및 3.34% 향상시키고 Synapse에서 HD를 각각 12.62 mm 및 2.48 mm 감소시켰습니다.
  • Bi-Level Routing Attention은 일반 전체 어텐션보다 계산 비용이 낮은 상태에서 장거리 모델링을 가능하게 합니다(복잡도는 O((HW)^(4/3))로 감소).
  • SCCSA 스킵 연결은 다운샘플링 중 손실된 공간 정보를 회복하고 차원 간 특성 상호작용을 강화합니다.
  • BRAU-Net++은 세 개의 공용 데이터셋(Synapse, ISIC-2018, CVC-ClinicDB)에서 강건한 일반화 성능을 보입니다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.