[论文解读] Segmentation of glioblastomas in early post-operative multi-modal MRI with deep neural networks
本研究提出基于深度神经网络的自动化方法,用于早期术后多模态MRI中残余胶质母细胞瘤的分割,Dice分数最高达61%,肿瘤存在分类的平衡准确率约为80%。该模型仅使用术后T1w-CE、T1w和FLAIR序列,在12家医院中实现良好泛化,性能与专家人类评分者相当,为手动肿瘤勾画提供了可重复的替代方案。
Extent of resection after surgery is one of the main prognostic factors for patients diagnosed with glioblastoma. To achieve this, accurate segmentation and classification of residual tumor from post-operative MR images is essential. The current standard method for estimating it is subject to high inter- and intra-rater variability, and an automated method for segmentation of residual tumor in early post-operative MRI could lead to a more accurate estimation of extent of resection. In this study, two state-of-the-art neural network architectures for pre-operative segmentation were trained for the task. The models were extensively validated on a multicenter dataset with nearly 1000 patients, from 12 hospitals in Europe and the United States. The best performance achieved was a 61\% Dice score, and the best classification performance was about 80\% balanced accuracy, with a demonstrated ability to generalize across hospitals. In addition, the segmentation performance of the best models was on par with human expert raters. The predicted segmentations can be used to accurately classify the patients into those with residual tumor, and those with gross total resection.
研究动机与目标
- 开发一种用于早期术后MRI中残余胶质母细胞瘤分割的自动化方法,以提高预后判断的准确性。
- 通过减少人类专家手动勾画中固有的组间和组内变异,降低这种变异的影响。
- 利用多中心数据集,在欧洲和美国多家医院中验证深度学习模型的泛化性能。
- 评估术前影像是否对实现高分割性能是必要的,或仅使用术后扫描是否已足够。
- 建立一种临床可部署、开源的解决方案,用于胶质母细胞瘤术后随访中的自动化肿瘤分割。
提出的方法
- 使用早期术后多模态MRI序列(T1w-CE、T1w、FLAIR)训练了两种最先进的深度神经网络架构,用于胶质母细胞瘤分割。
- 在来自欧洲和美国12家医院的956例患者的多中心数据集上对模型进行验证,采用按医院分层的训练/验证/测试集划分。
- 使用一个独立的医院作为测试集以评估泛化能力,并基于八名人类标注者的共识金标准评估组间变异。
- 使用Dice分数衡量分割性能,肿瘤存在分类则采用平衡准确率。
- 评估模型在有无术前MRI输入情况下的表现,以分析术前数据的贡献。
- 表现最佳的模型已发布于Raidionics平台,供公开访问及未来临床验证。

实验结果
研究问题
- RQ1深度神经网络能否在早期术后胶质母细胞瘤MRI中实现与人类专家评分者相当的分割性能?
- RQ2这些模型在不同临床机构和扫描协议下的泛化能力如何?
- RQ3术前影像对实现高精度残余肿瘤分割是否必不可少,还是仅使用术后序列已足够?
- RQ4模型的性能在多大程度上与人类专家之间的组间变异相当或更优?
- RQ5自动化分割能否可靠地将患者分类为全切或存在残余肿瘤?
主要发现
- 最佳模型在残余肿瘤分割中达到61%的Dice分数,与平均人类专家评分者的性能相当。
- 模型在分类患者是否存在残余肿瘤或全切方面,平衡准确率约为80%。
- 即使仅使用术后MRI进行训练和测试,性能依然稳健,无需依赖术前扫描。
- 模型优于新手标注者,且在与共识金标准对比时,性能与或超过个别专家标注。
- 人类专家之间的组间变异较高,而模型性能处于该范围内,支持其作为可靠替代方案的使用。
- 模型在12家医院中成功验证,表现出在多样化临床环境和成像协议下的强大泛化能力。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。