[论文解读] Image Quality Assessment for Magnetic Resonance Imaging
本研究对磁共振成像(MRI)图像质量评估(IQA)度量进行了迄今为止最大规模的评估,基于七名放射科医生提供的14,700个主观评分。研究识别出DISTS、HaarPSI、VSI和FID VGG16为表现最佳的度量,这些度量在信号噪声比(SNR)、对比噪声比(CNR)和伪影存在性等关键MRI质量标准上,与放射科医生的感知保持一致的高相关性。
Image quality assessment (IQA) algorithms aim to reproduce the human's perception of the image quality. The growing popularity of image enhancement, generation, and recovery models instigated the development of many methods to assess their performance. However, most IQA solutions are designed to predict image quality in the general domain, with the applicability to specific areas, such as medical imaging, remaining questionable. Moreover, the selection of these IQA metrics for a specific task typically involves intentionally induced distortions, such as manually added noise or artificial blurring; yet, the chosen metrics are then used to judge the output of real-life computer vision models. In this work, we aspire to fill these gaps by carrying out the most extensive IQA evaluation study for Magnetic Resonance Imaging (MRI) to date (14,700 subjective scores). We use outputs of neural network models trained to solve problems relevant to MRI, including image reconstruction in the scan acceleration, motion correction, and denoising. Our emphasis is on reflecting the radiologist's perception of the reconstructed images, gauging the most diagnostically influential criteria for the quality of MRI scans: signal-to-noise ratio, contrast-to-noise ratio, and the presence of artifacts. Seven trained radiologists assess these distorted images, with their verdicts then correlated with 35 different image quality metrics (full-reference, no-reference, and distribution-based metrics considered). The top performers -- DISTS, HaarPSI, VSI, and FID-VGG16 -- are found to be efficient across three proposed quality criteria, for all considered anatomies and the target tasks.
研究动机与目标
- 评估35种图像质量度量在预测放射科医生对MRI质量感知方面的表现。
- 评估现有IQA度量(专为自然图像设计)在MRI领域中的泛化能力,该领域具有独特的图像特征和诊断优先级。
- 识别能可靠反映临床相关MRI质量标准(如信噪比(SNR)、对比噪声比(CNR)和伪影存在性)的度量。
- 通过自一致性分析验证放射科医生标注的可靠性,并评估影响评分者间一致性的因素。
- 通过分析人类感知与度量预测之间的差异,为未来开发MRI专用、可解释的IQA度量提供指导。
提出的方法
- 从七名经验丰富的放射科医生处收集了14,700个主观质量评分,用于评估深度学习模型在多种解剖区域和任务中重建的成对MRI图像。
- 评估了35种IQA度量:13种全参考(FR)度量、12种无参考(NR)度量和10种基于分布(DB)的度量,包括FID、DISTS、VSI、HaarPSI、BRISQUE、PaQ-2-PiQ和MetaIQA。
- 采用成对图像比较方法评估放射科医生的感知,重点关注SNR、CNR和伪影水平作为关键诊断质量指标。
- 通过计算放射科医生评分与度量预测之间的Spearman等级相关系数,对不同解剖区域和任务中的度量性能进行排序。
- 通过重新标注部分图像对开展自一致性研究,以评估评分者间的一致性,并分析标注时间对一致性的影响。
- 通过比较MRI与自然图像数据集(TID2013、KADID-10k)上的度量表现,分析领域偏移效应,揭示不同度量在鲁棒性上的差异。
实验结果
研究问题
- RQ1在多种解剖区域和重建任务中,哪些图像质量评估度量与放射科医生对MRI质量的感知相关性最强?
- RQ2考虑到图像统计特性和诊断优先级的领域差异,专为自然图像设计的IQA度量在MRI上的表现如何?
- RQ3现代无参考和基于分布的IQA度量(如FID、DISTS)在在多大程度上能反映临床相关的MRI质量标准,如SNR、CNR和伪影存在性?
- RQ4哪些因素影响放射科医生标注的一致性?在大规模MRI IQA评估中,主观评分的可靠性如何?
- RQ5现有IQA度量是否可信赖地用于指导现实世界MRI应用中的模型选择,还是需要开发MRI专用、可解释的度量?
主要发现
- DISTS、HaarPSI、VSI和FID VGG16表现最为优异,其在所有评估的解剖区域和任务中均与放射科医生评分保持强且一致的相关性。
- 即使这些顶级度量最初是基于自然图像训练的,它们在MRI数据上仍保持高性能,表明对领域偏移具有鲁棒性。
- 基于分布的度量(如FID VGG16和DISTS)在MRI特定质量评估中优于传统FR和NR度量,尤其在捕捉与SNR和CNR相关的感知质量方面表现更优。
- 自一致性分析显示放射科医生之间存在中等到高度的一致性(加权Cohen’s Kappa值),且标注时间与一致性之间无显著相关性,表明专家判断具有稳定性。
- BRISQUE、PaQ-2-PiQ和MetaIQA等度量在MRI中表现较差,表明其从自然图像领域泛化能力有限。
- 本研究揭示了一个关键差距:尚无单一度量能完全捕捉所有临床相关的MRI质量方面,凸显了开发MRI专用、可解释的IQA模型的迫切需求。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。