[论文解读] 3D Tomographic Pattern Synthesis for Enhancing the Quantification of COVID-19
该论文提出了一种基于3D GAN的断层扫描模式生成框架,用于扩充有限的真实COVID-19 CT数据,从而提升肺部和异常区域分割的性能。通过结合位置先验和图像修复技术生成合成的COVID-19模式,该方法提高了分割准确性和定量分析的可靠性,使病灶包含率提升了6.02%,Percentage of Opacity(PO)计算的皮尔逊相关系数提高了2.82%。
The Coronavirus Disease (COVID-19) has affected 1.8 million people and resulted in more than 110,000 deaths as of April 12, 2020. Several studies have shown that tomographic patterns seen on chest Computed Tomography (CT), such as ground-glass opacities, consolidations, and crazy paving pattern, are correlated with the disease severity and progression. CT imaging can thus emerge as an important modality for the management of COVID-19 patients. AI-based solutions can be used to support CT based quantitative reporting and make reading efficient and reproducible if quantitative biomarkers, such as the Percentage of Opacity (PO), can be automatically computed. However, COVID-19 has posed unique challenges to the development of AI, specifically concerning the availability of appropriate image data and annotations at scale. In this paper, we propose to use synthetic datasets to augment an existing COVID-19 database to tackle these challenges. We train a Generative Adversarial Network (GAN) to inpaint COVID-19 related tomographic patterns on chest CTs from patients without infectious diseases. Additionally, we leverage location priors derived from manually labeled COVID-19 chest CTs patients to generate appropriate abnormality distributions. Synthetic data are used to improve both lung segmentation and segmentation of COVID-19 patterns by adding 20% of synthetic data to the real COVID-19 training data. We collected 2143 chest CTs, containing 327 COVID-19 positive cases, acquired from 12 sites across 7 countries. By testing on 100 COVID-19 positive and 100 control cases, we show that synthetic data can help improve both lung segmentation (+6.02% lesion inclusion rate) and abnormality segmentation (+2.78% dice coefficient), leading to an overall more accurate PO computation (+2.82% Pearson coefficient).
研究动机与目标
- 解决用于训练稳健AI模型的标注数据少、多样性不足且代表性差的COVID-19 CT影像数据问题。
- 克服由于疫情突发性带来的数据多样性、泛化能力及标注可扩展性方面的挑战。
- 提高如Percentage of Opacity(PO)和Pulmonary Involvement Score(PHO)等定量CT生物标志物的准确性和可重复性。
- 开发一种合成数据生成框架,以在3D胸部CT中保留临床相关的空间和纹理模式。
- 证明合成数据增强可提升肺部和异常区域分割的性能,使其接近人读之间的变异水平。
提出的方法
- 训练一个3D条件生成对抗网络(cGAN),利用单一异常标签在非感染CT扫描上执行COVID-19模式的图像修复。
- 将从人工标注的COVID-19 CT中提取的位置先验信息用于引导合成异常的空间分布。
- 使用3D形状生成算法来建模常见模式(如磨玻璃影、实变和碎石路征)的空间形态。
- 通过在真实COVID-19病例中加入20%的合成病例,对肺部分割和异常区域分割网络的训练数据进行增强。
- 在包含和不包含合成数据的情况下,对2D和3D U-Net基分割模型进行微调,以进行对比。
- 使用病灶包含率(LIR)、Dice相似系数(DSC)以及PO和PHO的皮尔逊相关系数(PCC)评估性能。
实验结果
研究问题
- RQ1基于GAN的3D断层扫描模式生成是否能有效扩充现实世界中有限的COVID-19 CT数据,从而提升分割性能?
- RQ2合成数据在多大程度上提升了对异常区域的检测能力,特别是那些在COVID-19中常见于外周肺区的区域?
- RQ3与真实值和人读之间变异相比,合成数据的引入如何影响Percentage of Opacity(PO)等定量生物标志物的准确性?
- RQ4使用真实标注中的位置先验是否能增强生成模式的解剖学合理性与临床相关性?
- RQ5该合成数据增强策略是否能在病灶分割与定量分析中达到与人类读者变异水平相当的性能?
主要发现
- 增加20%的合成数据使肺部分割的病灶包含率(LIR)提升了6.02%,表明对异常区域的覆盖更全面。
- 在使用合成数据的情况下,3D分割网络的Dice系数达到0.7064,相比无合成数据的基线(0.657)提升了2.78%。
- 在3D模型中,Percentage of Opacity(PO)的皮尔逊相关系数(PCC)从0.933提升至0.961,接近人读之间变异水平(PCC = 0.957)。
- 在3D模型中,Pulmonary Involvement Score(PHO)的PCC从0.9099提升至0.9387,表明与真实值的一致性显著增强。
- 在使用合成数据的3D模型中,PHO的DSC达到0.9387,与人读之间变异的DSC(0.7132 ± 0.1831)相当。
- PO相关系数的提升(PCC提高2.82%)表明,合成数据增强了定量CT生物标志物计算的可靠性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。