Skip to main content
QUICK REVIEW

[論文レビュー] Deep Learning for the Digital Pathologic Diagnosis of Cholangiocarcinoma and Hepatocellular Carcinoma: Evaluating the Impact of a Web-based Diagnostic Assistant

Bora Uyumazturk, Amirhossein Kiani|arXiv (Cornell University)|Nov 18, 2019
Cholangiocarcinoma and Gallbladder Cancer Studies参考文献 12被引用数 12
ひとこと要約

本研究では、全スライド画像における肝細胞癌(HCC)と胆管癌(CC)を区別するためのウェブベースのディープラーニング診断アシスタントの評価を行った。独立したテストセットにおいて84.2%の正確性を達成したが、病理医の診断正確性は向上せず、モデルの予測が病理医に顕著なバイアスをもたらした。正しければ性能が向上し、間違っていれば悪化する傾向にあり、AI支援病理診断におけるアンカリングバイアスのリスクを浮き彫りにした。

ABSTRACT

While artificial intelligence (AI) algorithms continue to rival human performance on a variety of clinical tasks, the question of how best to incorporate these algorithms into clinical workflows remains relatively unexplored. We investigated how AI can affect pathologist performance on the task of differentiating between two subtypes of primary liver cancer, hepatocellular carcinoma (HCC) and cholangiocarcinoma (CC). We developed an AI diagnostic assistant using a deep learning model and evaluated its effect on the diagnostic performance of eleven pathologists with varying levels of expertise. Our deep learning model achieved an accuracy of 0.885 on an internal validation set of 26 slides and an accuracy of 0.842 on an independent test set of 80 slides. Despite having high accuracy on a hold out test set, the diagnostic assistant did not significantly improve performance across pathologists (p-value: 0.184, OR: 1.287 (95% CI 0.886, 1.871)). Model correctness was observed to significantly bias the pathologist decisions. When the model was correct, assistance significantly improved accuracy across all pathologist experience levels and for all case difficulty levels (p-value: < 0.001, OR: 4.289 (95% CI 2.360, 7.794)). When the model was incorrect, assistance significantly decreased accuracy across all 11 pathologists and for all case difficulty levels (p-value < 0.001, OR: 0.253 (95% CI 0.126, 0.507)). Our results highlight the challenges of translating AI models to the clinical setting, especially for difficult subspecialty tasks such as tumor classification. In particular, they suggest that incorrect model predictions could strongly bias an expert's diagnosis, an important factor to consider when designing medical AI-assistance systems.

研究の動機と目的

  • ウェブベースのディープラーニング診断アシスタントが、HCCとCCを区別する際の病理医の診断正確性を向上させるかを評価すること。
  • モデルの予測の正しさが病理医の意思決定に与える影響、特に診断バイアスのリスクを調査すること。
  • 実臨床ワークフローのシミュレーションにおいて、経験水準の異なる病理医におけるアシスタントのパフォーマンスを評価すること。
  • サブスペシャリティ病理診断分野におけるAI意思決定支援ツールの実装可能性と安全性を検討すること。

提案手法

  • 70枚のH&E染色全スライド画像(HCC 35枚、CC 35枚)を用いて、The Cancer Genome AtlasのデータからDenseNet-121畳み込みニューラルネットワークをトレーニングした。
  • 腫瘍領域を対象に画像パッチを抽出し、モデルのトレーニングおよび検証に使用。性能評価は内部バリデーションセット(26枚)および独立した外部テストセット(80枚)で実施した。
  • クラウドデプロイされたウェブインターフェースにより、病理医が選択した画像パッチをアップロードし、リアルタイムでAIフィードバックを受け取る仕組みを実装し、臨床的第二意見ツールをシミュレートした。
  • 11名の病理医(トレーニー、非消化器専門医、消化器専門医、NOC病理医を含む)が、クロスオーバー設計により80枚のWSIを解釈し、半数の症例では支援あり、半数では支援なしで評価した。
  • 混合効果ロジスティック回帰モデルを用いて、支援の有無およびモデルの正しさが診断正確性に与える影響を評価し、有意性はWaldカイ二乗検定で検証した。
  • スライドレベルでのモデルパフォーマンスは、HCCまたはCCを分類するための0.5の確率閾値を用いて評価した。

実験結果

リサーチクエスチョン

  • RQ1ディープラーニングベースの診断アシスタントの統合により、病理医のHCCとCCの区別における診断正確性は向上するか?
  • RQ2AIモデルの予測の正しさは、病理医の診断パフォーマンスにどのように影響するか?
  • RQ3アシスタントの影響は、経験水準の異なる病理医においても一様か?
  • RQ4AIアシスタントの予測が誤りである場合、診断バイアス、特にアンカリングバイアスがどの程度発生するか?

主な発見

  • ディープラーニングモデルは、内部バリデーションセットで88.5%の正確性、独立した外部テストセット(80枚の全スライド画像)で84.2%の正確性を達成した。
  • 支援ありの全体的な病理医の診断正確性に有意な向上は認められなかった(p値:0.184、オッズ比:1.287、95%信頼区間 0.886–1.871)。
  • AIモデルが正しかった場合、病理医の正確性は顕著に向上した(p < 0.001、オッズ比:4.289、95%信頼区間 2.360–7.794)。
  • AIモデルが誤っていた場合、病理医の正確性は顕著に低下した(p < 0.001、オッズ比:0.253、95%信頼区間 0.126–0.507)。
  • モデル出力によるバイアス効果は、全病理医の経験水準および症例の難易度レベルにおいて一貫していた。
  • 本研究は、誤ったAI予測が即座に熟練病理医でさえ強く誤導する可能性があるアンカリングバイアスの深刻なリスクを明らかにした。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。