Skip to main content
QUICK REVIEW

[论文解读] Foundational Models in Medical Imaging: A Comprehensive Survey and Future Vision

Bobby Azad, Reza Azad|arXiv (Cornell University)|Oct 28, 2023
Radiomics and Machine Learning in Medical Imaging被引用 40
一句话总结

本综述评估医学影像中的基础模型,提出一个分类法,讨论训练、提示和多模态适应,并概述挑战与未来方向。

ABSTRACT

Foundation models, large-scale, pre-trained deep-learning models adapted to a wide range of downstream tasks have gained significant interest lately in various deep-learning problems undergoing a paradigm shift with the rise of these models. Trained on large-scale dataset to bridge the gap between different modalities, foundation models facilitate contextual reasoning, generalization, and prompt capabilities at test time. The predictions of these models can be adjusted for new tasks by augmenting the model input with task-specific hints called prompts without requiring extensive labeled data and retraining. Capitalizing on the advances in computer vision, medical imaging has also marked a growing interest in these models. To assist researchers in navigating this direction, this survey intends to provide a comprehensive overview of foundation models in the domain of medical imaging. Specifically, we initiate our exploration by providing an exposition of the fundamental concepts forming the basis of foundation models. Subsequently, we offer a methodical taxonomy of foundation models within the medical domain, proposing a classification system primarily structured around training strategies, while also incorporating additional facets such as application domains, imaging modalities, specific organs of interest, and the algorithms integral to these models. Furthermore, we emphasize the practical use case of some selected approaches and then discuss the opportunities, applications, and future directions of these large-scale pre-trained models, for analyzing medical images. In the same vein, we address the prevailing challenges and research pathways associated with foundational models in medical imaging. These encompass the areas of interpretability, data management, computational requirements, and the nuanced issue of contextual comprehension.

研究动机与目标

  • 提供医学影像中基础模型(FMs)及其核心概念的结构化概述。
  • 基于训练策略与模态引入医学FMs的分类法。
  • 分析跨影像模态与器官的应用以突出优点与局限。
  • 讨论可解释性、数据隐私、计算等方面的挑战,以及领域知识整合。
  • 为临床实践中的MFMs提出未来方向和研究路径。

提出的方法

  • 定义基础模型并总结其预训练目标(对比学习、生成、混合)。
  • 将医学FMs分为六类:VPM-Generalist、TPM-Hybrid、TPM-Contrastive、TPM-Generative、VPM-Adaptations、TPM-Conversational。
  • 回顾文本提示和视觉提示模型的代表性工作,重点关注医学影像任务。
  • 讨论提示工程、指令对齐,以及以语言模型为灵感的提示如何转化到视觉-语言医学任务。
  • 汇集开源实现并提供持续更新的参考(GitHub 计划)。
(a) Algorithms
(a) Algorithms

实验结果

研究问题

  • RQ1哪些基础概念和训练策略使医学基础模型成为可能?
  • RQ2医学影像FMs如何按提示和模态进行分类,以及每个类别的代表性方法是什么?
  • RQ3MFMs的关键临床机会、局限性与挑战(可解释性、隐私、计算、分布偏移)是什么?
  • RQ4哪些未来方向和悬而未决的研究问题将引导MFMs实现更广泛的临床影响?

主要发现

  • 基础模型在有限标注数据下实现多模态解释与上下文内适应。
  • 六组医学FMs分类涵盖文本提示和视觉提示,包括对话式和通用型变体。
  • 知名示例(如 MedCLIP、BiomedCLIP、MI-Zero、BioViL-T、Med-Flamingo)展示放射学、病理学等领域的零-shot、少样本和多模态能力。
  • 提示工程、指令对齐和领域知识整合是提升医学影像FM表现的核心。
  • 在医疗环境中,隐私保护数据使用和联邦学习被强调为关键优势。
  • 本综述强调持续更新和开源资源以实现快速传播与评估。
(b) Modalities
(b) Modalities

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。