[论文解读] Full-dose Whole-body PET Synthesis from Low-dose PET Using High-efficiency Denoising Diffusion Probabilistic Model: PET Consistency Model
本文提出PET-CM,一种高效去噪扩散概率模型,通过使用PET移位窗口视觉Transformer(PET-VIT)从低剂量输入合成全剂量全器官PET图像。该方法在去噪低剂量PET扫描方面实现了最先进图像质量,推理速度比以往扩散模型快12倍,展现出卓越的定量与临床性能。
Objective: Positron Emission Tomography (PET) has been a commonly used imaging modality in broad clinical applications. One of the most important tradeoffs in PET imaging is between image quality and radiation dose: high image quality comes with high radiation exposure. Improving image quality is desirable for all clinical applications while minimizing radiation exposure is needed to reduce risk to patients. Approach: We introduce PET Consistency Model (PET-CM), an efficient diffusion-based method for generating high-quality full-dose PET images from low-dose PET images. It employs a two-step process, adding Gaussian noise to full-dose PET images in the forward diffusion, and then denoising them using a PET Shifted-window Vision Transformer (PET-VIT) network in the reverse diffusion. The PET-VIT network learns a consistency function that enables direct denoising of Gaussian noise into clean full-dose PET images. PET-CM achieves state-of-the-art image quality while requiring significantly less computation time than other methods. Results: In experiments comparing eighth-dose to full-dose images, PET-CM demonstrated impressive performance with NMAE of 1.278+/-0.122%, PSNR of 33.783+/-0.824dB, SSIM of 0.964+/-0.009, NCC of 0.968+/-0.011, HRS of 4.543, and SUV Error of 0.255+/-0.318%, with an average generation time of 62 seconds per patient. This is a significant improvement compared to the state-of-the-art diffusion-based model with PET-CM reaching this result 12x faster. Similarly, in the quarter-dose to full-dose image experiments, PET-CM delivered competitive outcomes, achieving an NMAE of 0.973+/-0.066%, PSNR of 36.172+/-0.801dB, SSIM of 0.984+/-0.004, NCC of 0.990+/-0.005, HRS of 4.428, and SUV Error of 0.151+/-0.192% using the same generation process, which underlining its high quantitative and clinical precision in both denoising scenario.
研究动机与目标
- 为解决PET成像中辐射剂量与图像质量之间的临床权衡问题,通过从低剂量采集实现全剂量PET图像合成。
- 克服传统机器学习方法依赖手工设计特征、难以捕捉复杂解剖与纹理细节的局限性。
- 开发一种稳定、高保真度的图像合成方法,在定量指标与临床相关性方面均优于基于GAN的方法及以往扩散模型。
- 在基于扩散的PET图像合成中实现高计算效率,支持实际临床部署。
- 通过全面的定量指标与临床评估(包括SUV准确性与人工评分)验证方法性能。
提出的方法
- PET-CM采用两步扩散过程:前向扩散在全剂量PET图像上添加高斯噪声,反向扩散通过学习的一致性函数对图像进行去噪。
- 反向去噪过程通过概率流常微分方程(PF-ODE)建模,使用二阶Heun法求解器求解,实现稳定且精确的推理。
- 提出一种新型PET-VIT网络,基于Swin Transformer架构,学习一致性函数,将噪声潜在状态映射为干净的全剂量PET图像。
- PET-VIT在编码器与解码器模块之间集成残差连接与跳跃连接,以保持空间分辨率并改善特征传播。
- 在每个模块中注入使用正弦编码的时间步嵌入,以条件化去噪过程,使其依赖于扩散步骤。
- 通过预测前向扩散过程中添加的噪声,最小化去噪损失,实现反向过程的端到端学习。
实验结果
研究问题
- RQ1基于扩散的生成模型是否能在保持高计算效率的同时,实现从低剂量输入合成全剂量PET图像的最先进图像质量?
- RQ2所提出的PET-VIT架构在捕捉PET图像解剖与纹理细节方面,相较于标准CNN及其他注意力机制网络表现如何?
- RQ3与真实全剂量图像相比,PET-CM模型在保留定量指标(如标准化摄取值SUV)方面达到何种程度?
- RQ4在临床评估中,该模型表现如何,包括人工阅片者评分与与真实图像的结构相似性?
- RQ5所提出方法是否能显著缩短推理时间,同时不牺牲图像质量?
主要发现
- 在八分之一剂量至全剂量图像合成中,PET-CM实现NMAE为1.278±0.122%,PSNR为33.783±0.824 dB,SSIM为0.964±0.009,NCC为0.968±0.011,表现出优异的结构与强度保真度。
- 模型获得4.543的人工评分(HRS),表明在阅片者评估中其感知质量与全剂量扫描相当。
- SUV误差仅为0.255±0.318%,证实合成图像在肿瘤学应用中保留了关键的定量代谢准确性。
- 在四分之一剂量输入下,PET-CM实现NMAE为0.973±0.066%与PSNR为36.172±0.801 dB,表明在更低剂量水平下仍具强大鲁棒性。
- 模型平均仅需62秒即可生成一例全剂量PET图像,推理速度比最先进扩散模型快12倍。
- Swin Transformer架构与Heun法求解器的结合,实现了高保真、稳定且高效的推理,在训练稳定性与FID类指标方面优于GAN模型。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。