Skip to main content
QUICK REVIEW

[論文レビュー] The Impact of an Inter-rater Bias on Neural Network Training

Or Shwartzman, Harel Gazit|arXiv (Cornell University)|Jun 12, 2019
AI in cancer detection参考文献 11被引用数 4
ひとこと要約

本稿は、被験者間バイアス(被験者間での手動セグメンテーションアノテーションの差異)が、医療画像セグメンテーションの深層ニューラルネットワーク(DNN)学習に与える影響を調査する。ISBI 2015 MSチャレンジデータセットを用いて、1人の被験者からのアノテーションで学習したDNNが、別の被験者からのデータに対してテストされた場合、Diceスコアが低く、病変負荷の推定がバイアスを受けることを示した。これは、モデルの性能がランダムなばらつきではなく、被験者固有のバイアスに敏感であることを示している。

ABSTRACT

The problem of inter-rater variability is often discussed in the context of manual labeling of medical images. It is assumed to be bypassed by automatic model-based approaches for image segmentation which are considered `objective', providing single, deterministic solutions. However, the emergence of data-driven approaches such as Deep Neural Networks (DNNs) and their application to supervised semantic segmentation - brought this issue of raters' disagreement back to the front-stage. In this paper, we highlight the issue of inter-rater bias as opposed to random inter-observer variability and demonstrate its influence on DNN training, leading to different segmentation results for the same input images. In fact, lower Dice scores are calculated if training and test segmentations are of different raters. Moreover, we demonstrate that inter-rater bias in the training examples is amplified when considering the segmentation predictions for the test data. We support our findings by showing that a classifier-DNN trained to distinguish between raters based on their manual annotations performs better when the automatic segmentation predictions rather than the raters' annotations were tested. For this study, we used the ISBI 2015 Multiple Sclerosis (MS) challenge dataset, which includes annotations by two raters with different levels of expertise. The results obtained allow us to underline a worrisome clinical implication of a DNN bias induced by an inter-rater bias during training. Specially, we show that the differences in MS-lesion load estimates increase when the volume calculations are done based on the DNNs' segmentation predictions instead of the manual annotations used for training.

研究の動機と目的

  • 被験者間バイアス(ランダムなばらつきとは異なる)が、深層学習ベースの医療画像セグメンテーションに与える影響を調査すること。
  • トレーニングと推論の段階で被験者によるアノテーションの差がDNN性能に与える影響、特にDiceスコアと病変負荷推定の観点から評価すること。
  • DNNがトレーニングデータに存在する被験者固有のバイアスを学習し、それを予測にまで拡大するかどうかを評価すること。
  • このようなバイアスの臨床的意味、特に複数の硬化症病変の定量的MRI解析における影響を検討すること。

提案手法

  • ISBI 2015 マルチプルスクラーズチャレンジデータセット(異なる専門性を持つ2名の被験者によるアノテーションを含む)を用いた。
  • 1人の被験者からのアノテーションでDNNをトレーニングし、別の被験者からのテストデータで評価することで、被験者バイアスの影響を分離した。
  • 予測と異なる被験者からの正解アノテーションとの間のセグメンテーション被り具合を測定するためにDice係数を用いた。
  • 被験者を識別する分類器DNNを別にトレーニングし、その性能をDNN予測と元の手動アノテーションの両方でテストすることで、バイアスの拡散を評価した。
  • DNN予測から得られる病変負荷推定値と手動アノテーションとの比較を通じて、定量的指標におけるバイアスの拡大を評価した。

実験結果

リサーチクエスチョン

  • RQ1手動アノテーションにおける被験者間バイアスは、医療画像のセマンティックセグメンテーションにおけるDNNの性能にどのように影響するか?
  • RQ2テストアノテーションがトレーニングアノテーションとは異なる被験者からのものである場合、Diceスコアは低下するか?
  • RQ3DNNはトレーニングデータに存在する被騟能者固有のバイアスをどれほど学習し、予測にまで拡大するか?
  • RQ4被験者を識別するようにトレーニングした分類器が、DNN予測に対して元の手動アノテーションよりも高い精度を示す場合、バイアスの拡大が示唆されるか?
  • RQ5DNN予測を用いる場合、手動アノテーションと比較して、被験者間バイアスは定量的病変負荷推定にどのように影響するか?

主な発見

  • トレーニングとテストのアノテーションが異なる被験者からの場合、Diceスコアが低く観察された。これは、被験者間バイアスによる性能低下を示している。
  • 被験者を識別するDNN分類器は、DNN予測に対して手動アノテーションよりも高い正確性を達成した。これは、モデルが被験者固有のバイアスを拡大していることを示唆している。
  • トレーニングデータにおける被験者間バイアスは、DNNのセグメンテーション予測において拡大され、被験者間の病変負荷推定値の差が拡大した。
  • DNN予測から得られる定量的病変負荷推定値は、手動アノテーションよりも被験者間でばらつきが大きかった。これは、臨床的に懸念されるバイアスの拡大を示している。
  • 結果として、DNNが被験者バイアスに影響されないというわけではない。むしろ、決定論的モデルでトレーニングされた場合でさえ、被験者間の差を拡大・強調する可能性があることが示された。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。