[Paper Review] Segmentation of glioblastomas in early post-operative multi-modal MRI with deep neural networks
This study proposes deep neural networks for automated segmentation of residual glioblastoma in early post-operative multi-modal MRI, achieving Dice scores of up to 61% and balanced accuracy of ~80% for tumor presence classification. The models generalize across 12 hospitals using only post-operative T1w-CE, T1w, and FLAIR sequences, performing on par with expert human raters and offering a reproducible alternative to manual tumor delineation.
Extent of resection after surgery is one of the main prognostic factors for patients diagnosed with glioblastoma. To achieve this, accurate segmentation and classification of residual tumor from post-operative MR images is essential. The current standard method for estimating it is subject to high inter- and intra-rater variability, and an automated method for segmentation of residual tumor in early post-operative MRI could lead to a more accurate estimation of extent of resection. In this study, two state-of-the-art neural network architectures for pre-operative segmentation were trained for the task. The models were extensively validated on a multicenter dataset with nearly 1000 patients, from 12 hospitals in Europe and the United States. The best performance achieved was a 61\% Dice score, and the best classification performance was about 80\% balanced accuracy, with a demonstrated ability to generalize across hospitals. In addition, the segmentation performance of the best models was on par with human expert raters. The predicted segmentations can be used to accurately classify the patients into those with residual tumor, and those with gross total resection.
Motivation & Objective
- To develop an automated method for residual glioblastoma segmentation in early post-operative MRI to improve prognostic accuracy.
- To reduce inter- and intra-rater variability inherent in manual tumor delineation by human experts.
- To validate the generalization performance of deep learning models across multiple European and U.S. hospitals using a multicenter dataset.
- To assess whether pre-operative imaging is necessary for high segmentation performance or if post-operative scans alone suffice.
- To establish a clinically deployable, open-source solution for automated tumor segmentation in glioblastoma surgery follow-up.
Proposed method
- Two state-of-the-art deep neural network architectures were trained for glioblastoma segmentation using early post-operative multi-modal MRI sequences (T1w-CE, T1w, FLAIR).
- The models were validated on a multicenter dataset of 956 patients from 12 hospitals across Europe and the U.S., with hospital-stratified train/val/test splits.
- A hold-out hospital was used as a test set to evaluate generalization, with inter-rater variability assessed against consensus ground truth from eight human annotators.
- Segmentation performance was measured using the Dice score, while tumor presence classification used balanced accuracy.
- The models were evaluated both with and without pre-operative MRI inputs to assess the contribution of pre-operative data.
- The best-performing models were released in the Raidionics environment for open access and future clinical validation.

Experimental results
Research questions
- RQ1Can deep neural networks achieve segmentation performance comparable to human expert raters in early post-operative glioblastoma MRI?
- RQ2How well do these models generalize across diverse clinical sites and scanner protocols?
- RQ3Is pre-operative imaging necessary for high-accuracy residual tumor segmentation, or can post-operative sequences alone suffice?
- RQ4To what extent does the model’s performance match or exceed inter-rater variability among human experts?
- RQ5Can automated segmentation reliably classify patients as having gross total resection or residual tumor?
Key findings
- The best model achieved a Dice score of 61% for residual tumor segmentation, matching the performance of average human expert raters.
- The model demonstrated a balanced accuracy of approximately 80% in classifying patients as having residual tumor or gross total resection.
- Performance remained robust even when trained and tested on post-operative MRI alone, without requiring pre-operative scans.
- The model outperformed novice annotators and matched or exceeded individual expert annotations, even when evaluated against a consensus ground truth.
- Inter-rater variability among human experts was high, and the model’s performance fell within this range, supporting its use as a reliable alternative.
- The models were successfully validated across 12 hospitals, demonstrating strong generalization across diverse clinical settings and imaging protocols.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.