Skip to main content
QUICK REVIEW

[論文レビュー] Performance assessment of the deep learning technologies in grading glaucoma severity

Zhen Yi, Lei Wang|arXiv (Cornell University)|Oct 31, 2018
Retinal Imaging and Analysis被引用数 11
ひとこと要約

本研究では、色眼底画像を用いて緑内障の重症度を評価するための8つのディーブラーニングモデルを評価し、全体的および局所的領域の注目領域(ROI)における性能を比較した。特に微調整されたモデル、特にDenseNetは、全体的ROIで最大85.29%の二次加重カッパを達成し、全視野眼底画像が局所的視神経乳頭領域よりも豊富な診断的情報を含んでいることを示した。

ABSTRACT

Objective: To validate and compare the performance of eight available deep learning architectures in grading the severity of glaucoma based on color fundus images. Materials and Methods: We retrospectively collected a dataset of 5978 fundus images and their glaucoma severities were annotated by the consensus of two experienced ophthalmologists. We preprocessed the images to generate global and local regions of interest (ROIs), namely the global field-of-view images and the local disc region images. We then divided the generated images into three independent sub-groups for training, validation, and testing purposes. With the datasets, eight convolutional neural networks (CNNs) (i.e., VGG16, VGG19, ResNet, DenseNet, InceptionV3, InceptionResNet, Xception, and NASNetMobile) were trained separately to grade glaucoma severity, and validated quantitatively using the area under the receiver operating characteristic (ROC) curve and the quadratic kappa score. Results: The CNNs, except VGG16 and VGG19, achieved average kappa scores of 80.36% and 78.22% when trained from scratch on global and local ROIs, and 85.29% and 82.72% when fine-tuned using the pre-trained weights, respectively. VGG16 and VGG19 achieved reasonable accuracy when trained from scratch, but they failed when using pre-trained weights for global and local ROIs. Among these CNNs, the DenseNet had the highest classification accuracy (i.e., 75.50%) based on pre-trained weights when using global ROIs, as compared to 65.50% when using local ROIs. Conclusion: The experiments demonstrated the feasibility of the deep learning technology in grading glaucoma severity. In particular, global field-of-view images contain relatively richer information that may be critical for glaucoma assessment, suggesting that we should use the entire field-of-view of a fundus image for training a deep learning network.

研究の動機と目的

  • 色眼底画像から得た8つのディーブラーニングアーキテクチャの、緑内障重症度分類における性能を評価・比較すること。
  • 全体的視野画像と局所的視神経乳頭領域のどちらがより優れた診断的性能を示すかを評価すること。
  • 事前学習済み重みを用いた微調整と、ランダム初期化からの学習の両方のアプローチがモデルの精度に与える影響を特定すること。
  • 緑内障重症度分類に最も効果的なディーブラーニングアーキテクチャを同定すること。
  • 臨床的眼底画像を用いてディーブラーニングが、緑内障重症度を信頼性高く分類可能であることを検証すること。

提案手法

  • 2名の専門的眼科医による緑内障重症度のアノテーションがなされた5,978枚の眼底画像を、後向きに収集した。
  • 全体的視野および局所的視神経乳頭領域の注目領域(ROI)パッチを抽出するための画像前処理を実施した。
  • 各ROIタイプについて、独立した訓練・検証・テストサブセットにデータセットを分割した。
  • VGG16, VGG19, ResNet, DenseNet, InceptionV3, InceptionResNet, Xception, NASNetMobileの8つの事前学習済みおよびランダム初期化済みの畳み込みニューラルネットワーク(CNN)を、全体的および局所的ROIで学習させた。
  • 受信者操作特性曲線(ROC)下の面積と二次加重カッパスコアを用いて、モデルの性能を評価した。
  • 学習戦略(ランダム初期化からの学習対比して微調整)の違いによる性能の比較を実施した。

実験結果

リサーチクエスチョン

  • RQ1どのディーブラーニングアーキテクチャが眼底画像からの緑内障重症度分類において最も優れた性能を示すか?
  • RQ2全体的視野画像を用いることで、局所的視神経乳頭領域を用いる場合よりも分類性能が向上するか?
  • RQ3事前学習済み重みを用いた微調整は、ランダム初期化からの学習に比べてモデル性能にどのように影響するか?
  • RQ4VGGベースのモデルは、事前学習済み重みを用いた場合でも一貫した性能を示すのか、それとも特定の状況で失敗するのか?
  • RQ5グローバルROIとローカルROIの両方に対して適用した場合、モデル間で顕著な性能差が生じるか?

主な発見

  • DenseNetは、全体的ROIで事前学習済み重みを用いた場合に75.50%の最高分類精度を達成したのに対し、局所的ROIでは65.50%であった。
  • 微調整されたモデルは、ランダム初期化からの学習よりも平均カッパスコアが高かった(全体的ROIで85.29%、局所的ROIで82.72%)。
  • VGG16およびVGG19は、ランダム初期化からの学習では妥当な性能を示したが、事前学習済み重みを用いた場合、全体的および局所的ROIの両方で失敗した。
  • 全体的視野画像は、すべてのモデルにおいて一貫して優れた性能を示し、緑内障重症度評価に豊富な診断的情報を含んでいることが示された。
  • 最も優れた性能を示したモデル(DenseNet、全体的ROI、微調整)は、二次加重カッパスコア85.29%を達成した。
  • 本研究は、標準的な眼底画像を用いてディーブラーニングが自動的に緑内障重症度を分類可能であることを確認した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。