Skip to main content
QUICK REVIEW

[论文解读] Dataset Growth in Medical Image Analysis Research

Yuval E. Landau, Nahum Kiryati|arXiv (Cornell University)|Aug 21, 2019
Radiomics and Machine Learning in Medical Imaging参考文献 14被引用 4
一句话总结

本研究通过分析2011–2018年MICCAI会议论文,探讨了医学影像分析中数据集规模的趋势,揭示了在MRI、CT和fMRI模态中,受试者数量持续呈现指数级增长。作者发现,七年间数据集的中位数规模增长了3至10倍,年均增长率达21%–31%,表明同行评审在无形中将更大规模的数据集作为事实上的标准。

ABSTRACT

Medical image analysis studies usually require medical image datasets for training, testing and validation of algorithms. The need is underscored by the deep learning revolution and the dominance of machine learning in recent medical image analysis research. Nevertheless, due to ethical and legal constraints, commercial conflicts and the dependence on busy medical professionals, medical image analysis researchers have been described as "data starved". Due to the lack of objective criteria for sufficiency of dataset size, the research community implicitly sets ad-hoc standards by means of the peer review process. We hypothesize that peer review requires researchers to report the use of ever-increasing datasets as one condition for acceptance of their work to reputable publication venues. To test this hypothesis, we scanned the proceedings of the eminent MICCAI (Medical Image Computing and Computer-Assisted Intervention) conferences from 2011 to 2018. From a total of 2136 articles, we focused on 907 papers involving human datasets of MRI (Magnetic Resonance Imaging), CT (Computed Tomography) and fMRI (functional MRI) images. For each modality, for each of the years 2011-2018 we calculated the average, geometric mean and median number of human subjects used in that year's MICCAI articles. The results corroborate the dataset growth hypothesis. Specifically, the annual median dataset size in MICCAI articles has grown roughly 3-10 times from 2011 to 2018, depending on the imaging modality. Statistical analysis further supports the dataset growth hypothesis and reveals exponential growth of the geometric mean dataset size, with annual growth of about 21% for MRI, 24% for CT and 31% for fMRI. In slight analogy to Moore's law, the results can provide guidance about trends in the expectations of the medical image analysis community regarding dataset size.

研究动机与目标

  • 调查由于隐性同行评审标准,医学影像分析中的数据集规模要求是否随时间推移而增加。
  • 量化同行评审医学影像分析研究中数据集规模的增长趋势,尤其是高影响力会议中的表现。
  • 评估医学影像分析领域是否已形成对更大数据集的隐性期望,类似于摩尔定律。
  • 分析不同成像模态(MRI、CT和fMRI)随时间的数据集规模增长模式。
  • 提供实证证据,表明尽管存在伦理和后勤限制,医学影像分析研究中的数据集规模期望仍在持续增加。

提出的方法

  • 作者收集并分析了907篇来自2011–2018年MICCAI会议的论文,这些论文使用了来自MRI、CT和fMRI模态的人体医学影像数据集。
  • 针对每种模态和年份,计算了所有相关论文中受试者数量的平均值、几何平均值和中位数。
  • 应用统计分析评估数据集规模增长趋势,包括对几何平均值拟合指数增长模型。
  • 本研究使用MICCAI会议论文集作为高影响力医学影像分析研究期刊的代表性样本。
  • 分析严格聚焦于使用人体受试者的研究,以确保与临床和伦理限制的相关性。
  • 作者比较了不同模态之间的增长率,以评估数据集规模演变的差异。

实验结果

研究问题

  • RQ12011年至2018年期间,同行评审研究中使用的医学影像数据集规模是否显著增加?
  • RQ2不同成像模态(MRI、CT、fMRI)在数据集规模增长方面是否表现出不同的趋势?
  • RQ3医学影像分析研究中,数据集规模是否呈现出指数级增长的证据?
  • RQ4同行评审过程在多大程度上隐性地强制要求更大的数据集?
  • RQ5在不同年份和模态之间,中位数、平均值和几何平均值数据集规模如何比较?

主要发现

  • 2011年至2018年期间,MICCAI论文中的数据集中位数规模根据成像模态不同,增长了3至10倍。
  • 几何平均数据集规模以每年21%的速率呈指数增长(MRI),24%(CT),31%(fMRI)。
  • 统计分析证实,所有三种模态的数据集规模均呈现显著且持续的上升趋势。
  • 尽管在数据收集方面持续存在伦理、法律和后勤限制,数据集规模的增长仍被观察到。
  • 研究结果表明,同行评审已隐性地确立了更大数据集的事实标准,即使没有正式标准。
  • 观察到的增长模式具有指数特性,类似于摩尔定律,表明数据需求存在自我强化的趋势。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。