[論文レビュー] Explainable Convolutional Neural Networks for Retinal Fundus Classification and Cutting-Edge Segmentation Models for Retinal Blood Vessels from Fundus Images
本論文は、XAI法と組み合わせた8つの事前学習CNNを視網膜 fundus 画像分類で比較し、2つのデータセットFIVESとDRIVEを用いて視網膜血管の高度なセグメンテーションモデルを評価する。分類で高い精度(例:ResNet101 94.17%)とセグメンテーションで強い性能(例:Swin-Unet 86.19% Mean Pixel Accuracy)を報告する。
Our research focuses on the critical field of early diagnosis of disease by examining retinal blood vessels in fundus images. While automatic segmentation of retinal blood vessels holds promise for early detection, accurate analysis remains challenging due to the limitations of existing methods, which often lack discrimination power and are susceptible to influences from pathological regions. Our research in fundus image analysis advances deep learning-based classification using eight pre-trained CNN models. To enhance interpretability, we utilize Explainable AI techniques such as Grad-CAM, Grad-CAM++, Score-CAM, Faster Score-CAM, and Layer CAM. These techniques illuminate the decision-making processes of the models, fostering transparency and trust in their predictions. Expanding our exploration, we investigate ten models, including TransUNet with ResNet backbones, Attention U-Net with DenseNet and ResNet backbones, and Swin-UNET. Incorporating diverse architectures such as ResNet50V2, ResNet101V2, ResNet152V2, and DenseNet121 among others, this comprehensive study deepens our insights into attention mechanisms for enhanced fundus image analysis. Among the evaluated models for fundus image classification, ResNet101 emerged with the highest accuracy, achieving an impressive 94.17%. On the other end of the spectrum, EfficientNetB0 exhibited the lowest accuracy among the models, achieving a score of 88.33%. Furthermore, in the domain of fundus image segmentation, Swin-Unet demonstrated a Mean Pixel Accuracy of 86.19%, showcasing its effectiveness in accurately delineating regions of interest within fundus images. Conversely, Attention U-Net with DenseNet201 backbone exhibited the lowest Mean Pixel Accuracy among the evaluated models, achieving a score of 75.87%.
研究の動機と目的
- 視覚 fundus 画像分類のための8つの事前学習CNNモデルを複数のXAI手法と組み合わせて解釈性を高めることを評価する。
- 高度なセグメンテーションアーキテクチャ(Attention U-Net系、TransUNet、Swin-Unet)を視網膜血管の境界決定に適用する。
- 標準的な指標を用いて公開データセット上で分類およびセグメンテーションタスクの性能を比較する。
- 信頼できる fundus 画像解析のために最も効果的なモデルと explainability 手法を特定する。
提案手法
- 8つの事前学習CNNをベンチマーク:ResNet101、DenseNet169、Xception、InceptionV3、DenseNet121、InceptionResNetV2、ResNet50、EfficientNetB0。
- Grad-CAM、Grad-CAM++、Score-CAM、Faster Score-CAM、Layer CAM などのXAI手法を適用して予測を解釈する。
- DenseNet系・ResNet系などのバックボーンを用いたAttention U-Net系、TransUNet、Swin-Unetを用いたセグメンテーションを評価する。
- データセット:FIVES で分類、DRIVE および FIVES でセグメンテーション;Accuracy、Precision、Recall、F1、Jaccard、Log Loss(分類)および IoU、Dice、Mean Pixel Accuracy、Mean Modified Hausdorff Distance、Mean Surface Dice(セグメンテーション)などの指標を報告する。
- 性能比較:分類での最高精度(例:ResNet101 94.17%)およびセグメンテーション(例:Swin-Unet 86.19% Mean Pixel Accuracy)を報告する。
- 解釈性の結果とアーキテクチャ間のモデル挙動について議論する。

実験結果
リサーチクエスチョン
- RQ1どの事前学習CNNが多様なXAI解釈と組み合わせたとき、 fundus 画像分類で最高の精度を達成するか?
- RQ2高度なセグメンテーションモデル(Attention U-Net系、TransUNet、Swin-Unet)は、視網膜血管のセグメンテーションをベースラインと比較してどの程度の性能を示すか?
- RQ3異なるXAI手法が fundus 分類モデルの interpretability に与える影響はどの程度か?
- RQ4DRIVE および FIVES データセットで最良のセグメンテーション指標を達成するバックボーンの選択は?
主な発見
- ResNet101 は分類で最高の精度を 94.17% で達成した。
- EfficientNetB0 は分類精度が 88.33% で最も低かった。
- Swin-Unet はセグメンテーションで Mean Pixel Accuracy が 86.19% である。
- Attention U-Net の DenseNet201 バックボーンはセグメンテーションの Mean Pixel Accuracy が 75.87% で最も低かった。
- 本研究は Grad-CAM、Grad-CAM++、Score-CAM、Faster Score-CAM、Layer CAM を用いてモデル決定を明らかにする。

より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。