Skip to main content
QUICK REVIEW

[论文解读] WORD: A large scale dataset, benchmark and clinical applicable study for abdominal organ segmentation from CT image

Xiangde Luo, Wenjun Liao|arXiv (Cornell University)|Nov 3, 2021
Radiomics and Machine Learning in Medical Imaging被引用 11
一句话总结

本文提出了 WORD,一个大规模、临床标注的腹部CT数据集,包含150个体素的腹部CT影像,对16个器官进行了像素级分割和涂鸦标注,可作为全腹部器官分割深度学习模型的基准测试数据。研究揭示了当前最先进模型与临床医生在小而复杂的器官上的显著性能差距,凸显了在临床部署中亟需改进的、更高效标注且具备良好泛化能力的方法。

ABSTRACT

Whole abdominal organ segmentation is important in diagnosing abdomen lesions, radiotherapy, and follow-up. However, oncologists' delineating all abdominal organs from 3D volumes is time-consuming and very expensive. Deep learning-based medical image segmentation has shown the potential to reduce manual delineation efforts, but it still requires a large-scale fine annotated dataset for training, and there is a lack of large-scale datasets covering the whole abdomen region with accurate and detailed annotations for the whole abdominal organ segmentation. In this work, we establish a new large-scale extit{W}hole abdominal extit{OR}gan extit{D}ataset ( extit{WORD}) for algorithm research and clinical application development. This dataset contains 150 abdominal CT volumes (30495 slices). Each volume has 16 organs with fine pixel-level annotations and scribble-based sparse annotations, which may be the largest dataset with whole abdominal organ annotation. Several state-of-the-art segmentation methods are evaluated on this dataset. And we also invited three experienced oncologists to revise the model predictions to measure the gap between the deep learning method and oncologists. Afterwards, we investigate the inference-efficient learning on the WORD, as the high-resolution image requires large GPU memory and a long inference time in the test stage. We further evaluate the scribble-based annotation-efficient learning on this dataset, as the pixel-wise manual annotation is time-consuming and expensive. The work provided a new benchmark for the abdominal multi-organ segmentation task, and these experiments can serve as the baseline for future research and clinical application development.

研究动机与目标

  • 为解决缺乏大规模、高质量、全腹部CT数据集,且其多器官分割标注详尽的问题。
  • 利用真实临床数据,建立全腹部分割任务的标准化基准,用于评估深度学习模型的性能。
  • 通过将模型预测结果与经验丰富的肿瘤科医生的判断进行对比,探究深度学习模型在临床中的适用性。
  • 探索使用稀疏涂鸦标注的弱监督学习方法,以降低标注成本。
  • 评估模型在多样化数据集上的泛化能力,识别腹部分割任务中的领域差异问题。

提出的方法

  • 构建一个大规模真实临床数据集(WORD),包含150个体部CT影像,每个影像均对16个器官进行像素级标注和涂鸦标注。
  • 整合了30,495幅切片,其标注由放射科医生和肿瘤科医生进行精细且专家验证的标注。
  • 在WORD基准上评估12种最先进分割模型的性能,以评估其性能表现和临床可行性。
  • 组织一项用户研究,邀请三位初级肿瘤科医生对模型预测结果进行修订,量化模型输出与临床标准之间的差距。
  • 提出一种基于涂鸦的弱监督学习方法,通过最小化熵和类内强度方差,提升标注效率。
  • 通过对比在WORD、BTCV、TCIA以及外部LiTS数据集上的性能,评估模型的领域泛化能力,识别领域偏移问题。

实验结果

研究问题

  • RQ1在具有详细器官标注的大规模全腹部CT数据集上,当前最先进深度学习模型的性能如何?
  • RQ2在未经过修订的情况下,深度学习模型的预测结果在临床上可接受的程度如何,特别是针对小而复杂的器官?
  • RQ3基于涂鸦的弱监督学习在降低标注成本的同时,能否有效维持分割性能?
  • RQ4WORD数据集与其他公开数据集(如BTCV和TCIA)之间存在多大的领域偏移?这种偏移如何影响模型的泛化能力?
  • RQ5在真实临床工作流程中部署深度学习模型进行腹部多器官分割时,面临哪些关键挑战?

主要发现

  • 最先进模型在肝脏、脾脏和肾脏等大器官上表现优异,经肿瘤科医生轻微修订后具备临床可接受性。
  • 在胆囊、食管、胰腺和十二指肠等小而复杂的器官上,模型与肿瘤科医生之间存在显著性能差距,表明其临床准备度仍有限。
  • 基于涂鸦的学习方法相较于基线模型有所改进,但仍无法达到密集标注的性能水平,凸显了在高效标注学习方面仍有改进空间。
  • WORD数据集与其他公开数据集(如BTCV和TCIA)之间存在显著的领域差距,表明在某一数据集上训练的模型可能难以泛化到其他数据集。
  • WORD数据集是首个大规模、全腹部、多器官标注的临床数据集,具备全面的标注信息,为未来研究提供了公平的基准。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。