Skip to main content
QUICK REVIEW

[论文解读] Knowledge-based in silico models and dataset for the comparative evaluation of mammography AI for a range of breast characteristics, lesion conspicuities and doses

Elena Sizikova, Niloufar Saharkhiz|arXiv (Cornell University)|Oct 27, 2023
Radiomics and Machine Learning in Medical Imaging被引用 6
一句话总结

本文提出了一种基于物理的体外仿真流程,结合基于知识的数字乳腺模型与蒙特卡洛X射线仿真,生成用于评估人工智能在不同乳腺特征、病灶明显度及辐射剂量下性能的逼真合成乳腺X线摄影数据集。主要贡献是发布了M-SYNTH数据集——包含1,200个合成病例,涵盖不同的乳腺密度、肿块大小、肿块密度和曝光水平——结果表明,人工智能性能随乳腺密度增加而下降,且在较低辐射剂量下表现更优,验证了该方法作为可扩展、保护隐私的真人患者数据替代方案的有效性。

ABSTRACT

To generate evidence regarding the safety and efficacy of artificial intelligence (AI) enabled medical devices, AI models need to be evaluated on a diverse population of patient cases, some of which may not be readily available. We propose an evaluation approach for testing medical imaging AI models that relies on in silico imaging pipelines in which stochastic digital models of human anatomy (in object space) with and without pathology are imaged using a digital replica imaging acquisition system to generate realistic synthetic image datasets. Here, we release M-SYNTH, a dataset of cohorts with four breast fibroglandular density distributions imaged at different exposure levels using Monte Carlo x-ray simulations with the publicly available Virtual Imaging Clinical Trial for Regulatory Evaluation (VICTRE) toolkit. We utilize the synthetic dataset to analyze AI model performance and find that model performance decreases with increasing breast density and increases with higher mass density, as expected. As exposure levels decrease, AI model performance drops with the highest performance achieved at exposure levels lower than the nominal recommended dose for the breast type.

研究动机与目标

  • 为解决真实患者数据集的局限性,例如解剖变异有限、隐私限制以及病灶边界缺乏真实标签。
  • 开发一种可扩展、符合隐私合规要求的方法,用于在多样化的患者和成像参数下评估人工智能性能。
  • 构建一种基于仿真的框架,实现对人工智能模型在乳腺X线摄影中可控、可重复且伦理安全的测试。
  • 验证基于物理仿真的合成数据能否反映真实世界中人工智能性能的趋势,特别是关于乳腺密度和辐射剂量的影响。
  • 发布一个公开可用的数据集(M-SYNTH),以支持哺乳期人工智能模型的可重现性与对比性评估。

提出的方法

  • 该方法利用随机知识基(KB)的人体解剖模型生成1,200个具有不同纤维腺体密度、肿块大小和肿块密度的唯一数字乳腺体模。
  • 每个体模通过数字乳腺摄影系统的数字复制品进行成像,使用VICTRE工具包进行蒙特卡洛仿真,模拟X射线采集过程。
  • 成像流程在四个乳腺密度类别中控制曝光水平,支持剂量-反应分析。
  • 生成的合成图像用于通过AUC等指标评估人工智能模型性能,并对比不同物理和成像参数下的表现。
  • 该方法可完全控制对象和成像参数,包括肿块位置和范围,而这些在真实数据中常缺失。
  • 数据集已预先计算并通过GitHub发布,以支持离线评估并减轻用户计算负担。
Figure 1: Overview of the computational pipeline components for generating the M-SYNTH in silico dataset for medical imaging AI evaluation.
Figure 1: Overview of the computational pipeline components for generating the M-SYNTH in silico dataset for medical imaging AI evaluation.

实验结果

研究问题

  • RQ1在合成乳腺X线图像中,人工智能模型性能如何随乳腺密度增加而变化?
  • RQ2病灶明显度(大小和密度)如何影响在不同乳腺密度下的人工智能检测性能?
  • RQ3辐射剂量如何影响人工智能性能,且性能是否在低于标准临床剂量的水平达到峰值?
  • RQ4通过基于物理的仿真生成的合成数据能否反映真实世界中人工智能性能的趋势?
  • RQ5在图像统计特征和放射组学特征方面,模拟数据与真实患者数据的吻合程度如何?

主要发现

  • 人工智能模型性能(以AUC衡量)随乳腺密度增加而下降,随肿块大小和肿块密度增加而提高,符合预期。
  • 在较低辐射剂量下性能提升,且在每种乳腺类型的推荐剂量以下的曝光水平下,AUC达到最高值。
  • 合成图像在图像矩(均值、方差、偏度、峰度)方面与真实患者数据具有合理良好的一致性,尤其在全部四种乳腺密度中表现良好。
  • M-SYNTH数据集能够检测出由物理和成像参数引起的性能差异,证实该方法对关键变量具有敏感性。
  • 该方法成功验证了在真实患者数据(INBreast)上训练的人工智能模型在合成数据上的表现,支持该框架在监管评估中的实用性。
  • 本研究证明,基于物理的仿真可作为真实患者数据的可行、成本更低且保护隐私的替代方案,用于人工智能的对比评估。
Figure 6: Dose distribution (# of hist.) and percentages of optimal dose considered by breast density.
Figure 6: Dose distribution (# of hist.) and percentages of optimal dose considered by breast density.

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。