Skip to main content
QUICK REVIEW

[논문 리뷰] ColonFormer: An Efficient Transformer based Method for Colon Polyp Segmentation

Nguyen Thanh Duc, Thi-Oanh Nguyen|arXiv (Cornell University)|2022. 05. 17.
Colorectal Cancer Screening and Detection인용 수 4
한 줄 요약

ColonFormer는 다중 스케일 전역 특징 학습을 위한 경량 계층형 Transformer 인코더와 CNN 기반 계층형 디코더, 잔차 연결된 축 방향 attention을 사용한 정밀도 향상 모듈을 결합한 새로운 효율적인 Transformer 기반 인코더-디코더 네트워크로, 대장 폴립 세분화에 적용된다. 이는 다섯 개인 기준 데이터셋에서 최신 기술 수준의 성능을 달성하며, mDice 및 mIOU 지표에서 기존 방법들을 능가한다.

ABSTRACT

Identifying polyps is challenging for automatic analysis of endoscopic images in computer-aided clinical support systems. Models based on convolutional networks (CNN), transformers, and their combinations have been proposed to segment polyps with promising results. However, those approaches have limitations either in modeling the local appearance of the polyps only or lack of multi-level features for spatial dependency in the decoding process. This paper proposes a novel network, namely ColonFormer, to address these limitations. ColonFormer is an encoder-decoder architecture capable of modeling long-range semantic information at both encoder and decoder branches. The encoder is a lightweight architecture based on transformers for modeling global semantic relations at multi scales. The decoder is a hierarchical network structure designed for learning multi-level features to enrich feature representation. Besides, a refinement module is added with a new skip connection technique to refine the boundary of polyp objects in the global map for accurate segmentation. Extensive experiments have been conducted on five popular benchmark datasets for polyp segmentation, including Kvasir, CVC-Clinic DB, CVC-ColonDB, CVC-T, and ETIS-Larib. Experimental results show that our ColonFormer outperforms other state-of-the-art methods on all benchmark datasets.

연구 동기 및 목표

  • 기존 폴립 세분화 모델이 국소적 특징에만 집중하거나 다수준 공간적 의존성 모델링이 부족한 점을 해결하기 위해.
  • 낮은 대trast와 비정상적인 형태로 인해 경계가 모호한 폴립 영역의 세분화 정확도를 향상시키기 위해.
  • 임상 적용을 위한 효율적인 아키텍처를 개발하여 높은 성능와 낮은 계산 비용 사이의 균형을 이루기 위해.
  • 자기 주의를 통한 전역적 맥락 모델링과 계층적 특징 학습을 통합하여 다양한 영상 조건에서 강력한 폴립 탐지 능력을 확보하기 위해.

제안 방법

  • 인코더는 입력 이미지 전역에 걸쳐 다중 스케일 전역 의미 관계를 캡처하기 위해 경량 계층형 Transformer 아키텍처를 사용한다.
  • 디코더는 UPer 유사 피라미드 구조와 계층적 특징 융합을 활용하여 다중 수준 표현을 풍부하게 하고 공간적 이해도를 향상시킨다.
  • 새로운 정밀화 모듈을 도입하여 축 방향 주의와 잔차 연결을 통합하여 세분화 맵을 반복적으로 정밀화하고 경계 정확도를 향상시킨다.
  • 모델은 Transformer의 전역 모델링 능력과 CNN의 인덕티브 바이어스를 결합한 하이브리드 인코더-디코더 설계를 사용한다.
  • 정밀화 모듈은 거친 예측 결과를 네트워크를 거꾸로 전파하여 경계 오류를 수정하는 스킵 연결 메커니즘을 적용한다.
  • 아키텍처는 효율성 최적화가 이루어졌으며, 추론 실험 결과 UPer 디코더가 MLP 기반 디코더 대비 30% 이상 GFLOPs를 감소시켰다.

실험 결과

연구 질문

  • RQ1경량 Transformer 기반 인코더가 대장 폴립 세분화에 있어 장거리 의존성과 다중 스케일 특징을 효과적으로 모델링할 수 있는가?
  • RQ2계층적 CNN 기반 디코더의 통합이 표준 디코더 대비 특징 표현을 어떻게 향상시키는가?
  • RQ3제안된 축 주의와 잔차 연결을 갖춘 정밀화 모듈이 경계 세분화 정확도 향상에 얼마나 기여하는가?
  • RQ4다양한 폴립 세분화 기준 데이터셋에서 전체 ColonFormer 아키텍처가 최신 기술 수준의 모델 대비 성능과 효율성에서 어떻게 비교되는가?
  • RQ5다양한 백본 변종에서 모델 크기, 계산 비용, 세분화 정확도 사이의 최적의 트레이드오프는 무엇인가?

주요 결과

  • ColonFormer-S는 모든 다섯 개인 기준 데이터셋에서 평균 mDice 0.854를 기록하여 비교된 최신 기술 수준의 모든 방법들을 능가했다.
  • ColonFormer-L은 ETIS-Larib 데이터셋에서 mDice 0.801과 mIOU 0.722를 기록하여 이전 모델들보다 뚜렷이 뛰어난 성능을 보였다.
  • UPer 디코더는 MLP 디코더 대비 GFLOPs를 33% 감소시켰으며, 성능은 유사하게 유지되어 효율성의 타당성을 입증했다.
  • 정밀화 모듈은 Kvasir 및 ETIS-Larib 데이터셋에서 mDice를 최대 0.008 향상시켜 경계 정밀화 효과를 입증했다.
  • ColonFormer-L은 CVC-ClinicDB에서 mDice 0.932와 mIOU 0.884를 기록하여, SegFormer-B3-Uper-ARA 대비 각각 0.010과 0.009 높은 성능을 달성했다.
  • 추론 실험 결과, UPer 디코더와 정밀화된 축 주의 모듈의 조합이 최고의 전체 성능을 내며 최소한의 계산 오버헤드를 기록했다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.