Skip to main content
QUICK REVIEW

[논문 리뷰] MedSegDiff-V2: Diffusion based Medical Image Segmentation with Transformer

Junde Wu, Ji, Wei|arXiv (Cornell University)|2023. 01. 19.
Radiomics and Machine Learning in Medical Imaging인용 수 14
한 줄 요약

MedSegDiff-V2는 Transformer 기반 확산 프레임워크를 의료 영상 분할에 도입하고, Uncertain Spatial Attention이 있는 Anchor Condition과 Semantic Conditioning을 위한 Spectrum-Space Transformer (SS-Former)를 활용하여 다중 모달리티에서 20개의 분할 작업에 대해 최첨단 성능을 달성합니다.

ABSTRACT

The Diffusion Probabilistic Model (DPM) has recently gained popularity in the field of computer vision, thanks to its image generation applications, such as Imagen, Latent Diffusion Models, and Stable Diffusion, which have demonstrated impressive capabilities and sparked much discussion within the community. Recent investigations have further unveiled the utility of DPM in the domain of medical image analysis, as underscored by the commendable performance exhibited by the medical image segmentation model across various tasks. Although these models were originally underpinned by a UNet architecture, there exists a potential avenue for enhancing their performance through the integration of vision transformer mechanisms. However, we discovered that simply combining these two models resulted in subpar performance. To effectively integrate these two cutting-edge techniques for the Medical image segmentation, we propose a novel Transformer-based Diffusion framework, called MedSegDiff-V2. We verify its effectiveness on 20 medical image segmentation tasks with different image modalities. Through comprehensive evaluation, our approach demonstrates superiority over prior state-of-the-art (SOTA) methodologies. Code is released at https://github.com/KidsWithTokens/MedSegDiff

연구 동기 및 목표

  • UNet 백본 대비 분할 품질을 개선하기 위해 변환기(transformer)와 확산 기반 의학 영상 분할의 통합을 동기 부여한다.
  • 확산을 안정시키고 의미론적 상호작용을 강화하기 위해 Anchor Condition과 Semantic Condition의 두 가지 조건화 전략을 제안한다.
  • 주파수 도메인에서 노이즈와 의미 임베딩을 연결하기 위한 SS-Former를 개발한다.
  • 시퀀스 시간 스텝 전반에 걸쳐 확산 노이즈를 의미 특성에 맞추기 위해 적응형 Neural Band-pass Filter (NBP-Filter)를 도입한다.
  • 다양한 모달리티에 걸친 20개의 기관 분할 작업에서 SOTA 성능을 시연한다.

제안 방법

  • 두 개의 조건화 스트림: Anchor Condition은 Uncertain Spatial Attention (U-SA)을 통해 확산 인코더에 해독된 분할 특징을 주입하여 확산 분산을 줄인다.
  • Semantic Condition은 Spectrum-Space Transformer (SS-Former)와 주파수 공간의 Neural Band-pass Filter (NBP-Filter)를 사용하여 의미론적 분할 정보를 확산 임베딩에 삽입한다.
  • 확산 백본은 원시 이미지 특징에 조건화된 UNet 기반의 역과정으로 노이즈 제거 확산 확률 모델 (DPM)을 따른다.
  • U-SA는 가우시안 커널로 특징을 평활화하고 원래 특징과 최대값을 취하며 공간 주의(attention)와 유사한 1x1 컨브 변조를 적용하여 앵커 특징을 융합한다.
  • SS-Former는 조건과 확산 특징을 푸리에 공간으로 옮겨 교차 주의와 유사한 모듈로 정보를 교환하고, 확산 시간 스텝에 조건화된 NBP-Filter를 사용하여 스펙트럼을 정렬한다.
  • 훈련은 노이즈 예측 손실에 스케줄링된 조건화 감독을 포함한 앵커 손실(소프트 Dice + 교차 엔트로피)을 사용한다.
Figure 1: An illustration of MedSegDiff-V2, which starts from (a) an overview of the pipeline, and continues with zoomed-in diagrams of individual Models, including (b) SS-Former, and (c) NBP-Filter.
Figure 1: An illustration of MedSegDiff-V2, which starts from (a) an overview of the pipeline, and continues with zoomed-in diagrams of individual Models, including (b) SS-Former, and (c) NBP-Filter.

실험 결과

연구 질문

  • RQ1변환기 기반 조건화를 확산 모델과 통합하면 UNet 기반 확산 방법을 넘는 의료 영상 분할 성능을 향상시킬 수 있는가?
  • RQ2Anchor Condition이 변환기 백본을 사용할 때 확산 분산을 줄이고 안정성을 향상시키는가?
  • RQ3SS-Former가 주파수 도메인에서 확산 노이즈 임베딩과 의미 조건화를 효과적으로 결합하여 더 나은 분할을 이끌어낼 수 있는가?
  • RQ4다양한 모달리티에서 U-SA와 SS-Former가 정확도, 다양성, 수렴에 미치는 영향은 무엇인가?

주요 결과

  • MedSegDiff-V2는 5개 모달리티에 걸친 20개 기관 분할 작업에서 최첨단 성능을 달성합니다.
  • U-SA를 포함한 Anchor Condition은 일반 확산 성능을 크게 향상시키고 더 견고한 시작점을 제공합니다.
  • 특히 NBP-Filter와 결합될 때 SS-Former를 이용한 의미 조건은 노이즈와 의미 임베딩을 정렬하여 상당한 향상을 제공합니다.
  • 모델은 수렴에 필요한 앙상블 반복 수가 적고, Gflops가 낮아진 더 높은 Dice/IoU 지표를 제공하며 효율성이 향상됩니다.
  • 소거 실험은 Anchor Conditioning과 SS-Former의 분할 품질 향상 효과를 확인합니다.
Figure 2: The visual comparison with SOTA segmentation models on BTCV.
Figure 2: The visual comparison with SOTA segmentation models on BTCV.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.