Skip to main content
QUICK REVIEW

[Paper Review] Deep Learning for the Digital Pathologic Diagnosis of Cholangiocarcinoma and Hepatocellular Carcinoma: Evaluating the Impact of a Web-based Diagnostic Assistant

Bora Uyumazturk, Amirhossein Kiani|arXiv (Cornell University)|Nov 18, 2019
Cholangiocarcinoma and Gallbladder Cancer Studies12 references12 citations
TL;DR

This study evaluates a web-based deep learning diagnostic assistant for distinguishing hepatocellular carcinoma (HCC) and cholangiocarcinoma (CC) in whole-slide images. Despite achieving 84.2% accuracy on an independent test set, the assistant did not improve overall pathologist diagnostic accuracy, and model predictions significantly biased pathologists—improving performance when correct and worsening it when incorrect, highlighting risks of anchoring in AI-assisted pathology.

ABSTRACT

While artificial intelligence (AI) algorithms continue to rival human performance on a variety of clinical tasks, the question of how best to incorporate these algorithms into clinical workflows remains relatively unexplored. We investigated how AI can affect pathologist performance on the task of differentiating between two subtypes of primary liver cancer, hepatocellular carcinoma (HCC) and cholangiocarcinoma (CC). We developed an AI diagnostic assistant using a deep learning model and evaluated its effect on the diagnostic performance of eleven pathologists with varying levels of expertise. Our deep learning model achieved an accuracy of 0.885 on an internal validation set of 26 slides and an accuracy of 0.842 on an independent test set of 80 slides. Despite having high accuracy on a hold out test set, the diagnostic assistant did not significantly improve performance across pathologists (p-value: 0.184, OR: 1.287 (95% CI 0.886, 1.871)). Model correctness was observed to significantly bias the pathologist decisions. When the model was correct, assistance significantly improved accuracy across all pathologist experience levels and for all case difficulty levels (p-value: < 0.001, OR: 4.289 (95% CI 2.360, 7.794)). When the model was incorrect, assistance significantly decreased accuracy across all 11 pathologists and for all case difficulty levels (p-value < 0.001, OR: 0.253 (95% CI 0.126, 0.507)). Our results highlight the challenges of translating AI models to the clinical setting, especially for difficult subspecialty tasks such as tumor classification. In particular, they suggest that incorrect model predictions could strongly bias an expert's diagnosis, an important factor to consider when designing medical AI-assistance systems.

Motivation & Objective

  • To assess whether a web-based deep learning diagnostic assistant improves pathologist accuracy in differentiating HCC and CC.
  • To investigate the impact of model correctness on pathologist decision-making, particularly the risk of diagnostic bias.
  • To evaluate the assistant's performance across pathologists of varying expertise levels in a real-world clinical workflow simulation.
  • To examine the feasibility and safety of deploying AI decision support tools in subspecialty pathology settings.

Proposed method

  • A DenseNet-121 convolutional neural network was trained on 70 H&E-stained whole-slide images (35 HCC, 35 CC) from The Cancer Genome Atlas.
  • Image patches were extracted from tumor regions of interest and used to train and validate the model, with performance evaluated on an internal validation set (26 slides) and an independent external test set (80 slides).
  • A cloud-deployed web interface allowed pathologists to upload selected image patches for real-time AI feedback, simulating a clinical second-opinion tool.
  • Eleven pathologists (including trainees, non-GI specialists, GI specialists, and NOC pathologists) interpreted 80 WSI in a crossover design, with half the cases assisted and half unassisted.
  • Mixed-effects logistic regression models were used to assess the impact of assistance and model correctness on diagnostic accuracy, with significance tested via Wald Chi-square.
  • Model performance was evaluated at the slide level using a 0.5 probability threshold to classify HCC or CC.

Experimental results

Research questions

  • RQ1Does the integration of a deep learning-based diagnostic assistant improve the diagnostic accuracy of pathologists in distinguishing HCC from CC?
  • RQ2How does the correctness of the AI model's prediction affect pathologist diagnostic performance?
  • RQ3Does the impact of the assistant vary across pathologists of different experience levels?
  • RQ4To what extent does the AI assistant introduce diagnostic bias, particularly anchoring, when its predictions are incorrect?

Key findings

  • The deep learning model achieved 88.5% accuracy on the internal validation set and 84.2% accuracy on the independent external test set of 80 whole-slide images.
  • There was no statistically significant improvement in overall pathologist diagnostic accuracy with assistance (p-value: 0.184, OR: 1.287, 95% CI 0.886–1.871).
  • When the AI model was correct, pathologist accuracy improved significantly (p < 0.001, OR: 4.289, 95% CI 2.360–7.794).
  • When the AI model was incorrect, pathologist accuracy decreased significantly (p < 0.001, OR: 0.253, 95% CI 0.126–0.507).
  • The bias effect from model output was consistent across all pathologist experience levels and case difficulty levels.
  • The study reveals a critical risk of anchoring bias, where incorrect AI predictions strongly mislead even expert pathologists, undermining the safety of AI-assisted diagnosis.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.