Skip to main content
QUICK REVIEW

[論文レビュー] Robustness study of noisy annotation in deep learning based medical image segmentation

Shaode Yu, Erlei Zhang|arXiv (Cornell University)|Mar 10, 2020
Machine Learning and Data Classification参考文献 14被引用数 8
ひとこと要約

本研究は、頭頸部がん患者のCT画像を用いて、ノイズのあるアノテーション下での深層学習モデルの頬骨分類におけるロバスト性を調査する。ノイズあり・ノイズなしの下顎アノテーションの比率を変化させた訓練において、ノイズレベルが20%未満の場合には性能がわずかに低下するにとどまり、臨床現場におけるアノテーション誤差に対しても顕著なロバスト性を示している。

ABSTRACT

Partly due to the use of exhaustive-annotated data, deep networks have achieved impressive performance on medical image segmentation. Medical imaging data paired with noisy annotation are, however, ubiquitous, but little is known about the effect of noisy annotation on deep learning-based medical image segmentation. We studied the effects of noisy annotation in the context of mandible segmentation from CT images. First, 202 images of Head and Neck cancer patients were collected from our clinical database, where the organs-at-risk were annotated by one of 12 planning dosimetrists. The mandibles were roughly annotated as the planning avoiding structure. Then, mandible labels were checked and corrected by a physician to get clean annotations. At last, by varying the ratios of noisy labels in the training data, deep learning-based segmentation models were trained, one for each ratio. In general, a deep network trained with noisy labels had worse segmentation results than that trained with clean labels, and fewer noisy labels led to better segmentation. When using 20% or less noisy cases for training, no significant difference was found on the prediction performance between the models trained by noisy or clean. This study suggests that deep learning-based medical image segmentation is robust to noisy annotations to some extent. It also highlights the importance of labeling quality in deep learning

研究の動機と目的

  • ノイズのあるアノテーションが深層学習ベースの医療画像分類性能に与える影響を評価すること。
  • ノイズのあるアノテーションがモデル性能を著しく低下し始める閾値を定量化すること。
  • さまざまなレベルのアノテーションノイズを含むデータで訓練された深層学習モデルのロバスト性を評価すること。
  • 臨床画像における現実のアノテーション誤差に対する深層学習モデルの耐性について、実証的証拠を提供すること。

提案手法

  • 臨床データベースから12名の計画線量測定技師による初期アノテーションを含む202例のCT画像を収集した。
  • 医師による再評価と修正により、グランドトゥースのクリーンなアノテーションを取得した。
  • 訓練セットにおけるノイズあり・ノイズなしラベルの比率を系統的に変化させた(0%から100%まで)。
  • 同じアーキテクチャとトレーニングプロトコルを用いて、各ノイズ比率ごとに深層学習分類モデルを訓練した。
  • 保持されたテストセット上で、標準的な分類評価指標(例:Diceスコア)を用いてモデル性能を評価した。
  • 異なるノイズレベルで訓練されたモデル間の分類性能を比較し、ロバスト性を評価した。

実験結果

リサーチクエスチョン

  • RQ1ノイズありアノテーションの割合が増加するにつれて、深層学習モデルの分類性能にどのような影響を与えるか?
  • RQ2モデル性能が著しく低下し始めるノイズ比はどの程度か?
  • RQ3ノイズありアノテーションがモデル性能にほとんど影響を与えない閾値は存在するか?
  • RQ4ノイズありデータで訓練されたモデルの性能は、クリーンデータで訓練されたモデルと比べてどうか?

主な発見

  • ノイズラベルが20%以下で訓練されたモデルは、クリーンラベルで訓練されたモデルと比較して、統計的に有意な差が認められなかった。
  • 20%を超える訓練ラベルにノイズが含まれている場合にのみ、性能の低下が観察された。
  • 20%を超えるノイズラベルの割合が増加するにつれて、Diceスコアは段階的に低下した。
  • 本研究は、深層学習ベースの医療画像分類モデルが、中程度のアノテーションノイズに対して顕著なロバスト性を示すことを確認した。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。