Skip to main content
QUICK REVIEW

[논문 리뷰] ESFPNet: efficient deep learning architecture for real-time lesion segmentation in autofluorescence bronchoscopic video

Chang Qi, Danish Ahmad|arXiv (Cornell University)|2022. 07. 15.
Lung Cancer Diagnosis and Treatment인용 수 4
한 줄 요약

ESFPNet는 사전학습된 믹스 트랜스포머(MiT) 인코더와 효율적인 단계별 특징 피라미드(ESFP) 디코더를 활용하여 자가형광 내시경(AFB) 영상에서 기도 병변을 자동으로 분할하는 실시간 딥러닝 아키텍처이다. 평균 딱지 점수는 0.756이며, 1초당 27 프레임을 처리하여 실질적인 실시간 임상 적용이 가능하다.

ABSTRACT

Lung cancer tends to be detected at an advanced stage, resulting in a high patient mortality rate. Thus, much recent research has focused on early disease detection Bronchoscopy is the procedure of choice for an effective noninvasive way of detecting early manifestations (bronchial lesions) of lung cancer. In particular, autofluorescence bronchoscopy (AFB) discriminates the autofluorescence properties of normal (green) and diseased tissue (reddish brown) with different colors. Because recent studies show AFB's high sensitivity in searching lesions, it has become a potentially pivotal method in bronchoscopic airway exams. Unfortunately, manual inspection of AFB video is extremely tedious and error prone, while limited effort has been expended toward potentially more robust automatic AFB lesion analysis. We propose a real-time (processing throughput of 27 frames/sec) deep-learning architecture dubbed ESFPNet for accurate segmentation and robust detection of bronchial lesions in AFB video streams. The architecture features an encoder structure that exploits pretrained Mix Transformer (MiT) encoders and an efficient stage-wise feature pyramid (ESFP) decoder structure. Segmentation results from the AFB airway-exam videos of 20 lung cancer patients indicate that our approach gives a mean Dice index = 0.756 and an average Intersection of Union = 0.624, results that are superior to those generated by other recent architectures. Thus, ESFPNet gives the physician a potential tool for confident real-time lesion segmentation and detection during a live bronchoscopic airway exam. Moreover, our model shows promising potential applicability to other domains, as evidenced by its state-of-the-art (SOTA) performance on the CVC-ClinicDB, ETIS-LaribPolypDB datasets, and superior performance on the Kvasir, CVC-ColonDB datasets.

연구 동기 및 목표

  • 자가형광 내시경(AFB) 영상에서 병변 식별을 자동화하여 조기 폐암 검출의 필수적인 필요성을 충족시킨다.
  • 수동적인 AFB 영상 검토 방식의 한계를 극복한다. 이는 번거롭고 실수의 여지가 있으며 시간이 오래 걸린다.
  • 실시간으로 정확하고 강건한 딥러닝 모델을 개발하여 AFB 영상 스트림에서 기도 병변을 분할한다.
  • 기타 의료 영상 분야, 특히 내시경 외의 분야에서도 잘 일반화될 수 있도록 보장한다.
  • 분할 정확도를 희생시키지 않은 채 높은 추론 속도를 확보하여 실시간 내시경 중 사용이 가능하도록 한다.

제안 방법

  • 계층적 특징 추출을 위해 사전학습된 믹스 트랜스포머(MiT) 인코더를 활용한 U-Net 유사 인코더-디코더 아키텍처를 채택한다.
  • 다중 스케일 컨텍스트 간의 특징 집합을 향상시키기 위해 효율적인 단계별 특징 피라미드(ESFP) 디코더를 설계한다.
  • 다중 디코더 단계에서 특징 학습과 분할 정확도를 향상시키기 위해 깊은 감독을 통합한다.
  • 경량이며 파rameter 효율적인 구성 요소를 활용하여 높은 추론 속도(1초당 27 프레임)를 유지하면서도 성능을 유지를 한다.
  • 20명의 폐암 환자로부터 확보한 208개의 병변 영상 프레임으로 구성된 제한된 AFB 데이터셋에서 모델을 종합적으로 훈련한다.
  • 전이 학습 및 도메인 적응 기법을 적용하여 외부의 풀피프 분할 데이터셋에서의 일반화 능력을 향상시킨다.

실험 결과

연구 질문

  • RQ1딥러닝 모델이 자가형광 내시경 영상에서 기도 병변을 실시간으로 정확하게 분할할 수 있는가?
  • RQ2제안된 ESFPNet 아키텍처는 최신 기술(SOTA) 모델 대비 분할 정확도와 추론 속도 측면에서 어떻게 비교되는가?
  • RQ3ESFPNet 모델은 대肠 풍증 검출과 같은 다른 의료 영상 분할 작업으로까지 일반화될 수 있는가?
  • RQ4단계별 특징 피라미드 디코더는 기존 피라미드 구조에 비해 특징 표현을 효과적으로 향상시키는가?
  • RQ5제한된 임상 AFB 영상 데이터로부터 훈련된 모델이 높은 성능을 유지할 수 있는가?

주요 결과

  • ESFPNet는 AFB 폐암 환자 데이터셋에서 평균 딱지 점수 0.756과 평균 교차율(mIoU) 0.624를 기록하여 다른 최근 아키텍처를 능가했다.
  • 모델는 1초당 27 프레임을 처리하여 실시간 내시경 절차 중 적용이 가능했다.
  • CVC-ClinicDB 데이터셋에서 ESFPNet-L은 mDice 0.949와 mIoU 0.907을 기록하여 비교된 모든 모델 중 1위를 차지했다.
  • Kvasir-SEG 데이터셋에서 ESFPNet-L은 mDice 0.931과 mIoU 0.887을 기록하여 테스트된 모든 모델 중 2위를 기록했다.
  • 일반화성 테스트에서 ESFPNet-L은 ETIS-LaribPolypDB에서 mDice 0.827과 mIoU 0.752를 기록했으며, CVC-ColonDB에서는 mDice 0.823과 mIoU 0.741을 기록했다.
  • 모델는 다양한 데이터셋에서 뛰어난 성능을 보이며 의료 영상 분할 작업에서 높은 일반화 능력과 학습 능력을 입증했다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.