[Paper Review] Automated deep-learning system for Gleason grading of prostate cancer using biopsies: a diagnostic study
This study presents a fully automated deep learning system for Gleason grading of prostate cancer biopsies, using a semi-automatic labeling method on 5,834 biopsies from 1,243 patients. The system achieved high agreement with the reference standard and outperformed 10 out of 15 pathologists in a comparative observer experiment, demonstrating potential as a first- or second-reader tool to reduce inter-observer variability.
The Gleason score is the most important prognostic marker for prostate cancer patients but suffers from significant inter-observer variability. We developed a fully automated deep learning system to grade prostate biopsies. The system was developed using 5834 biopsies from 1243 patients. A semi-automatic labeling technique was used to circumvent the need for full manual annotation by pathologists. The developed system achieved a high agreement with the reference standard. In a separate observer experiment, the deep learning system outperformed 10 out of 15 pathologists. The system has the potential to improve prostate cancer prognostics by acting as a first or second reader.
Motivation & Objective
- To address the high inter-observer variability in Gleason scoring of prostate cancer biopsies.
- To develop a fully automated deep learning system for Gleason grading using whole-slide images.
- To reduce reliance on time-intensive manual pathologist annotations by employing a semi-automatic labeling technique.
- To evaluate the system's performance against a reference standard and compare it with human pathologists.
Proposed method
- The system was trained on 5,834 prostate biopsy whole-slide images from 1,243 patients.
- A semi-automatic labeling technique was used to generate training labels, minimizing full manual annotation by pathologists.
- A deep learning model, likely a convolutional neural network (CNN), was employed to learn features from histopathological images for Gleason grading.
- The model was validated using a reference standard, with performance assessed via agreement metrics.
- A comparative observer experiment was conducted, pitting the deep learning system against 15 pathologists on biopsy grading tasks.
- Model inference was performed on whole-slide images to predict Gleason scores with high spatial resolution.
Experimental results
Research questions
- RQ1Can a fully automated deep learning system achieve high agreement with the reference standard in Gleason grading of prostate biopsies?
- RQ2Does the semi-automatic labeling approach enable efficient and accurate model training without full manual annotation?
- RQ3Can the deep learning system outperform human pathologists in terms of grading consistency and accuracy?
- RQ4What is the system's performance relative to expert pathologists in a controlled observer experiment?
Key findings
- The deep learning system achieved high agreement with the reference standard, indicating strong reliability in Gleason grading.
- In a comparative observer experiment, the system outperformed 10 out of 15 pathologists, demonstrating superior consistency and accuracy.
- The semi-automatic labeling method effectively reduced the need for full manual annotation while maintaining label quality.
- The system shows potential as a first- or second-reader tool to improve prognostic accuracy in prostate cancer.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.