[论文解读] CLIP in Medical Imaging: A Survey
本综述分析了如何将 CLIP 适用于医学影像,详细介绍了精细化预训练方法与 CLIP 驱动的应用、挑战、数据集及未来方向。
Contrastive Language-Image Pre-training (CLIP), a simple yet effective pre-training paradigm, successfully introduces text supervision to vision models. It has shown promising results across various tasks due to its generalizability and interpretability. The use of CLIP has recently gained increasing interest in the medical imaging domain, serving as a pre-training paradigm for image-text alignment, or a critical component in diverse clinical tasks. With the aim of facilitating a deeper understanding of this promising direction, this survey offers an in-depth exploration of the CLIP within the domain of medical imaging, regarding both refined CLIP pre-training and CLIP-driven applications. In this paper, we (1) first start with a brief introduction to the fundamentals of CLIP methodology; (2) then investigate the adaptation of CLIP pre-training in the medical imaging domain, focusing on how to optimize CLIP given characteristics of medical images and reports; (3) further explore practical utilization of CLIP pre-trained models in various tasks, including classification, dense prediction, and cross-modal tasks; and (4) finally discuss existing limitations of CLIP in the context of medical imaging, and propose forward-looking directions to address the demands of medical imaging domain. Studies featuring technical and practical value are both investigated. We expect this survey will provide researchers with a holistic understanding of the CLIP paradigm and its potential implications. The project page of this survey can also be found on https://github.com/zhaozh10/Awesome-CLIP-in-Medical-Imaging.
研究动机与目标
- 提供 CLIP 概念及变体的全面概览。
- 分析 CLIP 预训练如何适应医学影像与报告。
- 总结基于 CLIP 的医学影像在各任务中的应用。
- 讨论挑战并提出医疗 CLIP 的未来研究方向。
提出的方法
- 给出与 CLIP 相关的医学影像研究的分类法。
- 描述对比学习预训练目标与零-shot 泛化方程(CLIP 中的方程(1)-(4))。
- 总结多尺度对比方法,如 GLoRIA 和 LoVT,以及它们对全球性 CLIP 的改进。
- 对医疗 CLIP 预训练中的数据高效和知识增强策略进行分类。
- 评审公开可用的医学图像文本数据集及相关的 CLIP 模型。

实验结果
研究问题
- RQ1如何将 CLIP 预训练适应医学影像和报告的特征?
- RQ2在医学数据中实现多尺度图文对齐的有效策略有哪些?
- RQ3数据高效性和知识引入如何提升医疗 CLIP 的性能?
- RQ4哪些任务和数据集能够展示 CLIP 驱动的医学影像能力?
主要发现
- CLIP 的图像-文本预训练可以扩展到医学影像领域,使零样本域识别和跨模态任务成为可能。
- 多尺度对比方法(如 GLoRIA、LoVT)提升了局部文本与局部图像的对齐,超越全局级别的 CLIP,有助于分割与检测。
- 数据高效策略(相关性驱动对比、句子/章节级提示、知识提示)可缓解小型医学数据集问题并提高鲁棒性。
- 一系列数据集(如 ROCO、MedICaT、PMC-OA、MIMIC-CXR、PadChest)支持医学图像文本研究,并在该领域中实现预训练的 CLIP 模型。
- 如 GLIP、CLIPSeg、CRIS 等变体将 CLIP 扩展到检测与分割,推动医学应用,如病变定位和句子定位。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。