Skip to main content
QUICK REVIEW

[論文レビュー] Quantitative and Qualitative Evaluation of Explainable Deep Learning Methods for Ophthalmic Diagnosis

Amitojdeep Singh, Janarthanam Jothi Balaji|arXiv (Cornell University)|Sep 26, 2020
Retinal Imaging and Analysis被引用数 5
ひとこと要約

本研究では、眼科臨床医のフィードバックを用いて、網膜光学干渉断層撮影(OCT)診断における深層学習モデルの13種類の説明可能なAI手法を評価した。Deep Taylorアトリビューションが最も高いスコア(中央値3.85/5)を記録し、ガイドドバックプロパゲーションやSHAPを上回り、眼科分野における臨床的採用のためのモデルの解釈可能性の重要性を示した。

ABSTRACT

Background: The lack of explanations for the decisions made by algorithms such as deep learning has hampered their acceptance by the clinical community despite highly accurate results on multiple problems. Recently, attribution methods have emerged for explaining deep learning models, and they have been tested on medical imaging problems. The performance of attribution methods is compared on standard machine learning datasets and not on medical images. In this study, we perform a comparative analysis to determine the most suitable explainability method for retinal OCT diagnosis. Methods: A commonly used deep learning model known as Inception v3 was trained to diagnose 3 retinal diseases - choroidal neovascularization (CNV), diabetic macular edema (DME), and drusen. The explanations from 13 different attribution methods were rated by a panel of 14 clinicians for clinical significance. Feedback was obtained from the clinicians regarding the current and future scope of such methods. Results: An attribution method based on a Taylor series expansion, called Deep Taylor was rated the highest by clinicians with a median rating of 3.85/5. It was followed by two other attribution methods, Guided backpropagation and SHAP (SHapley Additive exPlanations). Conclusion: Explanations of deep learning models can make them more transparent for clinical diagnosis. This study compared different explanations methods in the context of retinal OCT diagnosis and found that the best performing method may not be the one considered best for other deep learning tasks. Overall, there was a high degree of acceptance from the clinicians surveyed in the study. Keywords: explainable AI, deep learning, machine learning, image processing, Optical coherence tomography, retina, Diabetic macular edema, Choroidal Neovascularization, Drusen

研究の動機と目的

  • 実世界の専門家フィードバックを用いて、眼科診断における説明可能なAI手法の臨床的妥当性を評価すること。
  • 網膜OCT画像における深層学習モデルの最も効果的なアトリビューション手法を特定すること。
  • 眼科医における説明可能性技術の実用的有用性と受容性を評価すること。
  • 臨床的文脈における13種類のアトリビューション手法の定性的および定量的性能を比較すること。
  • 医療画像応用に特化した解釈可能なAIシステムの今後の開発を支援すること。

提案手法

  • 3つの網膜疾患(脈絡膜新生血管化(CNV)、糖尿病性黄斑浮腫(DME)、ドゥルーゼン)を分類する目的で、事前学習済みInception v3モデルをOCTスキャンに適応して微調整した。
  • Deep Taylor、ガイドドバックプロパゲーション、SHAPを含む13種類のアトリビューション手法を用い、モデル予測のためのサリエンシーマップを生成した。
  • 14名の眼科医によるパネルが、各説明の臨床的意義を5段階スケールで評価した。主に解剖学的妥当性と診断的関連性に注目した。
  • 定量的評価として、臨床医の評価結果を統計的に分析し、アトリビューション手法の順位を付与した。
  • 臨床医からの定性的フィードバックを収集し、説明可能性ツールの現在および将来の臨床ワークフローにおける有用性を評価した。
  • 本研究では、実世界の臨床的妥当性を保証するため、制御された「臨床医が関与する」評価フレームワークを採用した。

実験結果

リサーチクエスチョン

  • RQ1網膜OCT診断における深層学習モデルの説明可能なAI手法の中で、どの手法が最も臨床的意味を持つ説明を生成するか?
  • RQ2臨床医は、異なるアトリビューション手法の解釈可能性と解剖学的妥当性をどのように評価しているか?
  • RQ3臨床医は、診断意思決定におけるモデルの説明に対してどれほど信頼し、価値を見出しているか?
  • RQ4医療画像処理と標準ベンチマークとを比較した場合、アトリビューション手法の性能に差が生じるか?
  • RQ5説明可能なAIの臨床的および将来の応用について、臨床医はどのように認識しているか?

主な発見

  • Deep Taylorアトリビューションは、臨床医による中央値スコア3.85/5を記録し、強い臨床的解釈可能性を示した。
  • ガイドドバックプロパゲーションとSHAPは、それぞれ第2位および第3位の好まれ方を示し、Deep Taylorに次ぐ高いがわずかに低い評価を得た。
  • 臨床医は説明可能性手法の全体的な受容性が高く、臨床現場における信頼性向上と採用促進の可能性を強調した。
  • 本研究では、最良の性能を示す手法が文脈によって異なることが判明し、モデルの解釈可能性はドメイン特化された環境で評価されるべきであることを示唆した。
  • 臨床医のフィードバックから、診断的信頼性に寄与する説明における解剖学的正確性と局在化の重要性が強調された。
  • 臨床医は、既知の病理的特徴に対応する網膜の関心領域を明確に強調する手法を明確に好んでいた。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。