Skip to main content
QUICK REVIEW

[Paper Review] Bridging the gap between prostate radiology and pathology through machine learning

Indrani Bhattacharya, David Lim|arXiv (Cornell University)|Dec 3, 2021
Prostate Cancer Diagnosis and Treatment51 references17 citations
TL;DR

This study proposes using deep learning-generated digital pathologist labels—derived from whole-mount histopathology images—as superior training labels for machine learning models to detect and localize prostate cancer on MRI. These labels outperform radiologist and pathologist-annotated labels by enabling more accurate, consistent, and generalizable detection of both aggressive and indolent cancer components across diverse model architectures and patient cohorts.

ABSTRACT

Prostate cancer is the second deadliest cancer for American men. While Magnetic Resonance Imaging (MRI) is increasingly used to guide targeted biopsies for prostate cancer diagnosis, its utility remains limited due to high rates of false positives and false negatives as well as low inter-reader agreements. Machine learning methods to detect and localize cancer on prostate MRI can help standardize radiologist interpretations. However, existing machine learning methods vary not only in model architecture, but also in the ground truth labeling strategies used for model training. In this study, we compare different labeling strategies, namely, pathology-confirmed radiologist labels, pathologist labels on whole-mount histopathology images, and lesion-level and pixel-level digital pathologist labels (previously validated deep learning algorithm on histopathology images to predict pixel-level Gleason patterns) on whole-mount histopathology images. We analyse the effects these labels have on the performance of the trained machine learning models. Our experiments show that (1) radiologist labels and models trained with them can miss cancers, or underestimate cancer extent, (2) digital pathologist labels and models trained with them have high concordance with pathologist labels, and (3) models trained with digital pathologist labels achieve the best performance in prostate cancer detection in two different cohorts with different disease distributions, irrespective of the model architecture used. Digital pathologist labels can reduce challenges associated with human annotations, including labor, time, inter- and intra-reader variability, and can help bridge the gap between prostate radiology and pathology by enabling the training of reliable machine learning models to detect and localize prostate cancer on MRI.

Motivation & Objective

  • Address the persistent gap between prostate radiology and pathology in cancer detection due to inconsistent and error-prone human annotations.
  • Overcome limitations of radiologist-annotated labels, including high false positive/negative rates and inter-reader variability.
  • Evaluate the impact of different labeling strategies—radiologist, pathologist, and digital pathologist labels—on machine learning model performance for prostate MRI interpretation.
  • Demonstrate that digital pathologist labels, generated via deep learning on histopathology, enable superior and generalizable model performance across diverse datasets and architectures.
  • Enable selective identification of aggressive versus indolent cancer components in mixed lesions, a capability unattainable with human-annotated labels.

Proposed method

  • Trained four deep learning models—SPCNet, U-Net, branched U-Net, and DeepLabv3+—on pre-operative prostate MRI using four distinct label types: radiologist-confirmed, pathologist, and digital pathologist labels.
  • Mapped all pathology-based labels (pathologist and digital pathologist) onto pre-operative MRI using an automated MRI-histopathology registration platform.
  • Utilized a previously validated deep learning algorithm to generate pixel-level Gleason pattern labels on whole-mount histopathology slides, serving as digital pathologist labels.
  • Evaluated model performance on two independent cohorts: 40 radical prostatectomy patients (with whole-mount histopathology) and 275 targeted biopsy patients.
  • Measured performance using lesion-level and pixel-level metrics including ROC-AUC, Dice coefficient, and lesion volume overlap.
  • Compared model generalization and accuracy across label types and model architectures to isolate the impact of labeling strategy.

Experimental results

Research questions

  • RQ1How do different labeling strategies—radiologist-confirmed, pathologist, and digital pathologist labels—affect the performance of deep learning models in detecting prostate cancer on MRI?
  • RQ2Can digital pathologist-generated labels, derived from automated Gleason pattern prediction on histopathology, serve as a reliable and scalable alternative to human-annotated labels?
  • RQ3Does the use of digital pathologist labels improve the detection of both aggressive and indolent cancer components in mixed lesions compared to human-annotated labels?
  • RQ4How does model performance vary across different deep learning architectures when trained with different label types?
  • RQ5To what extent do digital pathologist labels reduce inter- and intra-observer variability and improve generalization across diverse clinical cohorts?

Key findings

  • Digital pathologist labels showed near-perfect concordance with pathologist labels (lesion ROC-AUC: 0.97–1.00, Dice: 0.75–0.93), significantly outperforming radiologist labels.
  • Machine learning models trained with digital pathologist labels achieved the highest lesion detection performance in the radical prostatectomy cohort (aggressive lesion ROC-AUC: 0.91–0.94).
  • Models trained with digital pathologist labels matched or exceeded the performance of pathologist label-trained models in the targeted biopsy cohort (aggressive lesion ROC-AUC: 0.87–0.88).
  • Digital pathologist label-trained models uniquely enabled pixel-level discrimination of aggressive and indolent cancer components within mixed lesions, a capability not possible with any human-annotated label type.
  • The performance advantage of digital pathologist labels was consistent across all four deep learning architectures (SPCNet, U-Net, branched U-Net, DeepLabv3+), indicating the benefit is independent of model architecture.
  • Radiologist labels missed up to 25% of pathology-confirmed lesions and showed low Dice overlap (0.24–0.28), underscoring their limitations for training reliable models.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.