Skip to main content
QUICK REVIEW

[논문 리뷰] Explainable Convolutional Neural Networks for Retinal Fundus Classification and Cutting-Edge Segmentation Models for Retinal Blood Vessels from Fundus Images

Fatema Tuj Johora Faria, Mukaffi Bin Moin|arXiv (Cornell University)|2024. 05. 12.
Retinal Imaging and Analysis인용 수 5
한 줄 요약

본 논문은 XAI 방법을 결합한 여덟 가지 사전 학습 CNN을 안저 이미지 분류에 대해 비교하고, 망막 혈관을 분할하기 위해 여러 고급 분할 모델을 평가했으며, 두 데이터셋(FIVES 및 DRIVE)을 사용한다. 분류 정확도는 높고(S 예: ResNet101 94.17%), 분할 성능도 우수하다고 보고한다(Swin-Unet 86.19% Mean Pixel Accuracy와 같다).

ABSTRACT

Our research focuses on the critical field of early diagnosis of disease by examining retinal blood vessels in fundus images. While automatic segmentation of retinal blood vessels holds promise for early detection, accurate analysis remains challenging due to the limitations of existing methods, which often lack discrimination power and are susceptible to influences from pathological regions. Our research in fundus image analysis advances deep learning-based classification using eight pre-trained CNN models. To enhance interpretability, we utilize Explainable AI techniques such as Grad-CAM, Grad-CAM++, Score-CAM, Faster Score-CAM, and Layer CAM. These techniques illuminate the decision-making processes of the models, fostering transparency and trust in their predictions. Expanding our exploration, we investigate ten models, including TransUNet with ResNet backbones, Attention U-Net with DenseNet and ResNet backbones, and Swin-UNET. Incorporating diverse architectures such as ResNet50V2, ResNet101V2, ResNet152V2, and DenseNet121 among others, this comprehensive study deepens our insights into attention mechanisms for enhanced fundus image analysis. Among the evaluated models for fundus image classification, ResNet101 emerged with the highest accuracy, achieving an impressive 94.17%. On the other end of the spectrum, EfficientNetB0 exhibited the lowest accuracy among the models, achieving a score of 88.33%. Furthermore, in the domain of fundus image segmentation, Swin-Unet demonstrated a Mean Pixel Accuracy of 86.19%, showcasing its effectiveness in accurately delineating regions of interest within fundus images. Conversely, Attention U-Net with DenseNet201 backbone exhibited the lowest Mean Pixel Accuracy among the evaluated models, achieving a score of 75.87%.

연구 동기 및 목표

  • 여러 XAI 기법으로 안저 이미지 분류를 위한 여덟 개의 사전 학습 CNN 모델을 평가하여 해석가능성을 높인다.
  • 망막 혈관 구분을 위한 고급 분할 아키텍처(Attention U-Net 변형, TransUNet, Swin-Unet)를 평가한다.
  • 공개 데이터셋에서 표준 지표를 사용하여 분류 및 분할 작업 간의 성능을 비교한다.
  • 신뢰할 수 있는 안저 이미지 분석을 위한 가장 효과적인 모델과 설명가능성 방법을 식별한다.

제안 방법

  • 사전 학습된 CNN 여덟 가지를 벤치마크한다: ResNet101, DenseNet169, Xception, InceptionV3, DenseNet121, InceptionResNetV2, ResNet50, EfficientNetB0.
  • Prediction 해석을 위해 Grad-CAM, Grad-CAM++, Score-CAM, Faster Score-CAM, Layer CAM과 같은 XAI 기법을 적용한다.
  • DenseNet 및 ResNet 계열과 같은 백본을 사용하여 Attention U-Net 변형, TransUNet, Swin-Unet으로 분할을 평가한다.
  • 데이터셋: FIVES에서 분류, DRIVE 및 FIVES에서 분할; 정확도(Accuracy), 정밀도(Precision), 재현율(Recall), F1, Jaccard, 로그 손실(Log Loss) 등 분류 지표 및 IoU, Dice, Mean Pixel Accuracy, Mean Modified Hausdorff Distance, Mean Surface Dice 등 분할 지표를 보고한다.
  • 비교 성능: 분류에서의 최고 정확도(예: 94.17%의 ResNet101)와 분할에서의 최고 성능(Swin-Unet 86.19% Mean Pixel Accuracy)을 보고한다.
  • 다양한 아키텍처의 설명가능성 결과와 모델 동작을 논의한다.
(a) AMD
(a) AMD

실험 결과

연구 질문

  • RQ1여러 XAI 설명과 함께 사용할 때 어떤 사전 학습 CNN이 안저 이미지 분류에서 가장 높은 정확도를 제공하는가?
  • RQ2고급 분할 모델(Attention U-Net 변형, TransUNet, Swin-Unet)은 기저선과 비교하여 망막 혈관 분할에서 어떤 성능을 보이는가?
  • RQ3다양한 XAI 기법이 망막 분류 모델의 해석가능성에 어떤 영향을 미치는가?
  • RQ4DRIVE 및 FIVES 데이터셋에서 어떤 백본 선택이 최고의 분할 지표를 산출하는가?

주요 결과

  • ResNet101은 분류 정확도에서 가장 높게 94.17%를 달성했다.
  • EfficientNetB0은 분류 정확도가 가장 낮은 88.33%를 보였다.
  • Swin-Unet은 분할에서 Mean Pixel Accuracy가 86.19%를 달성했다.
  • Attention U-Net의 DenseNet201 백본은 분할 Mean Pixel Accuracy가 75.87%로 가장 낮았다.
  • 연구는 Grad-CAM, Grad-CAM++, Score-CAM, Faster Score-CAM, Layer CAM을 사용하여 모델 결정을 밝힌다.
(b) DR
(b) DR

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.