Skip to main content
QUICK REVIEW

[論文レビュー] A Closer Look at Domain Shift for Deep Learning in Histopathology

Karin Stacke, Gabriel Eilertsen|arXiv (Cornell University)|Sep 25, 2019
AI in cancer detection参考文献 23被引用数 47
ひとこと要約

この論文は、スキャナー間でのCNN腫瘍分類器におけるヒストopathologyのドメインシフトを研究し、表現シフト指標を提案し、CAMELYON17データでデータ拡張・正規化手法を評価します。

ABSTRACT

Domain shift is a significant problem in histopathology. There can be large differences in data characteristics of whole-slide images between medical centers and scanners, making generalization of deep learning to unseen data difficult. To gain a better understanding of the problem, we present a study on convolutional neural networks trained for tumor classification of H&E stained whole-slide images. We analyze how augmentation and normalization strategies affect performance and learned representations, and what features a trained model respond to. Most centrally, we present a novel measure for evaluating the distance between domains in the context of the learned representation of a particular model. This measure can reveal how sensitive a model is to domain variations, and can be used to detect new data that a model will have problems generalizing to. The results show how learning is heavily influenced by the preparation of training data, and that the latent representation used to do classification is sensitive to changes in data distribution, especially when training without augmentation or normalization.

研究の動機と目的

  • スキャナー間でのH&E 全スライド画像におけるCNNベースの腫瘍分類に対するドメインシフトの影響を理解する。
  • 拡張と正規化戦略が性能と学習表現に与える影響を評価する。
  • ドメイン由来の特徴表現の変化を定量化する表現シフト指標を導入する。

提案手法

  • 3つのスキャナからのCAMELYON17 H&E WSIパッチを用いてクロスドメインシフトを模擬する。
  • 拡張/正規化の有無で、2つの小規模CNNアーキテクチャ(Simple CNNとMini-GoogLeNet)を評価する。
  • テストデータにカラー拡張、染色正規化、およびCycleGANに基づくドメイン翻訳を適用する。
  • モデルが依存する特徴を解釈するために学習されたフィルタを可視化する。
  • 最終畳込み層の各フィルタ活性化間のWasserstein距離に基づく表現シフト指標を提案・計算する。

実験結果

リサーチクエスチョン

  • RQ11つのスキャナで訓練し、未知のスキャナでテストした場合のクロスデータセット一般化はどのように機能するか?
  • RQ2異なるデータ処理戦略の下でCNNが学習する特徴は何か、そしてそれがドメインシフトへの頑健性にどう影響するか?
  • RQ3表現シフト指標はクロスドメイン分類精度の低下を予測できるか?

主な発見

  • 拡張なしでは、スキャナー間一般化が大幅に低下する(平均損失約21.7パーセントポイント)。
  • カラー拡張はスキャナー間一般化を小さな低下に改善(約4.75p.p.)。
  • 染色正規化は未知のケースの一部で有効だが、特定のスキャナーでより大きな低下を招くことがある(約9.2p.p.)。
  • CycleGANは同一スキャナー内の性能で最良を示すが、スキャナー間の一般化にはあまり適さず、最大で約11.45p.p.の低下。
  • 特徴の可視化はデータ処理により学習表現が異なることを示し、カラー拡張はモデルが色をより完全に無視するようにする。
  • 表現シフトとパッチレベル精度の間には負の相関があり、より大きなシフトはドメイン間でより低い精度を伴う傾向がある。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。