Skip to main content
QUICK REVIEW

[논문 리뷰] Quantitative and Qualitative Evaluation of Explainable Deep Learning Methods for Ophthalmic Diagnosis

Amitojdeep Singh, Janarthanam Jothi Balaji|arXiv (Cornell University)|2020. 09. 26.
Retinal Imaging and Analysis인용 수 5
한 줄 요약

이 연구는 망막 광학 간섭 단층촬영(OCT) 진료에서 딥러닝 모델을 위한 13가지 해석 가능한 AI 방법을 임상의 피드백을 통해 평가한다. Deep Taylor 할당법이 가장 높은 점수(중앙값 3.85/5)를 기록했으며, Guided Backpropagation 및 SHAP를 능가하여 모델의 해석 가능성은 안과 분야에서의 임상적 도입에 있어 핵심 요소임을 입증한다.

ABSTRACT

Background: The lack of explanations for the decisions made by algorithms such as deep learning has hampered their acceptance by the clinical community despite highly accurate results on multiple problems. Recently, attribution methods have emerged for explaining deep learning models, and they have been tested on medical imaging problems. The performance of attribution methods is compared on standard machine learning datasets and not on medical images. In this study, we perform a comparative analysis to determine the most suitable explainability method for retinal OCT diagnosis. Methods: A commonly used deep learning model known as Inception v3 was trained to diagnose 3 retinal diseases - choroidal neovascularization (CNV), diabetic macular edema (DME), and drusen. The explanations from 13 different attribution methods were rated by a panel of 14 clinicians for clinical significance. Feedback was obtained from the clinicians regarding the current and future scope of such methods. Results: An attribution method based on a Taylor series expansion, called Deep Taylor was rated the highest by clinicians with a median rating of 3.85/5. It was followed by two other attribution methods, Guided backpropagation and SHAP (SHapley Additive exPlanations). Conclusion: Explanations of deep learning models can make them more transparent for clinical diagnosis. This study compared different explanations methods in the context of retinal OCT diagnosis and found that the best performing method may not be the one considered best for other deep learning tasks. Overall, there was a high degree of acceptance from the clinicians surveyed in the study. Keywords: explainable AI, deep learning, machine learning, image processing, Optical coherence tomography, retina, Diabetic macular edema, Choroidal Neovascularization, Drusen

연구 동기 및 목표

  • 실제 임상 전문가 피드백을 활용하여 안과 진료에서 해석 가능한 AI 방법의 임상적 관련성을 평가하기 위해.
  • 망막 OCT 영상에서 딥러닝 모델을 위한 가장 효과적인 할당 방법을 특정하기 위해.
  • 안과의사들 사이에서 해석 가능성 기법의 실용적 유용성과 수용도를 평가하기 위해.
  • 임상 환경에서 13가지 할당 방법의 정성적 및 정량적 성능을 비교하기 위해.
  • 의료 영상 적용 분야에 맞는 해석 가능한 AI 시스템의 향후 개발을 안내하기 위해.

제안 방법

  • 3가지 망막 질환(모낭성 혈관신생, 당뇨성 망막부종, 드루센)을 분류하기 위해 사전 학습된 Inception v3 모델을 OCT 스캔에 대해 미세조정하였다.
  • Deep Taylor, Guided Backpropagation, SHAP 등을 포함한 13가지 할당 방법을 적용하여 모델 예측의 중요도 지도를 생성하였다.
  • 14명의 안과의사로 구성된 전문가 패anel이 각 설명의 임상적 의의를 5점 척도로 평가하였으며, 해부학적 타당성과 진단 관련성에 중점을 두었다.
  • 정량적 평가는 임상의 평가 점수의 통계 분석을 통해 할당 방법의 순위를 매기기 위해 실시되었다.
  • 임상의의 정성적 피드백을 수집하여 설명 가능성 도구의 현재 및 향후 임상 워크플로우 내 유용성 평가를 수행하였다.
  • 실제 임상적 관련성을 확보하기 위해 통제된, 임상의가 참여하는 평가 프레임워크를 사용하였다.

실험 결과

연구 질문

  • RQ1망막 OCT 진료에서 딥러닝 모델에 대한 가장 임상적으로 의미 있는 설명을 제공하는 해석 가능한 AI 방법은 무엇인가?
  • RQ2임상의들은 다양한 할당 방법의 해석 가능성과 해부학적 타당성에 대해 어떻게 평가하는가?
  • RQ3임상의들은 모델 설명이 진단 결정 과정에서 얼마나 신뢰하고 가치 있다고 보는가?
  • RQ4의료 영상 환경에서 평가했을 때와 표준 벤치마크에서 평가했을 때 할당 방법 간 성능 차이가 존재하는가?
  • RQ5해석 가능한 AI의 임상적 및 향후 적용 가능성이 안과 분야에서 어떻게 인식되는가?

주요 결과

  • Deep Taylor 할당법은 임상의 평가에서 중앙값 3.85/5로 가장 높은 점수를 기록하여 강력한 임상적 해석 가능성 임을 시사한다.
  • Guided Backpropagation와 SHAP는 각각 두 번째와 세 번째로 선호되는 방법이었으며, Deep Taylor에 비해 약간 낮은 점수를 기록하였다.
  • 임상의들은 설명 가능성 방법에 대해 전반적으로 높은 수용도를 보였으며, 임상 실무에서 신뢰도 향상과 도입 촉진 잠재력을 강조하였다.
  • 연구 결과, 최고 성능을 내는 방법은 맥락에 따라 달라졌으며, 이는 모델의 해석 가능성은 도메인 특화 환경에서 평가되어야 한다는 것을 시사한다.
  • 임상의 피드백은 진단적 자신감을 높이기 위해 설명의 해부학적 정확성과 국소화 능력이 중요하다고 강조하였다.
  • 임상의들은 알려진 병리적 특징에 해당하는 망막 영역을 강조하는 방법을 명확히 선호하였다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.