[Paper Review] Foundational Models in Medical Imaging: A Comprehensive Survey and Future Vision
This survey reviews foundation models in medical imaging, proposes a taxonomy, discusses training, prompting, and multimodal adaptation, and outlines challenges and future directions.
Foundation models, large-scale, pre-trained deep-learning models adapted to a wide range of downstream tasks have gained significant interest lately in various deep-learning problems undergoing a paradigm shift with the rise of these models. Trained on large-scale dataset to bridge the gap between different modalities, foundation models facilitate contextual reasoning, generalization, and prompt capabilities at test time. The predictions of these models can be adjusted for new tasks by augmenting the model input with task-specific hints called prompts without requiring extensive labeled data and retraining. Capitalizing on the advances in computer vision, medical imaging has also marked a growing interest in these models. To assist researchers in navigating this direction, this survey intends to provide a comprehensive overview of foundation models in the domain of medical imaging. Specifically, we initiate our exploration by providing an exposition of the fundamental concepts forming the basis of foundation models. Subsequently, we offer a methodical taxonomy of foundation models within the medical domain, proposing a classification system primarily structured around training strategies, while also incorporating additional facets such as application domains, imaging modalities, specific organs of interest, and the algorithms integral to these models. Furthermore, we emphasize the practical use case of some selected approaches and then discuss the opportunities, applications, and future directions of these large-scale pre-trained models, for analyzing medical images. In the same vein, we address the prevailing challenges and research pathways associated with foundational models in medical imaging. These encompass the areas of interpretability, data management, computational requirements, and the nuanced issue of contextual comprehension.
Motivation & Objective
- Provide a structured overview of foundational models (FMs) in medical imaging and their core concepts.
- Introduce a taxonomy of medical FMs based on training strategies and modalities.
- Analyze applications across imaging modalities and organs to highlight strengths and limitations.
- Discuss challenges such as interpretability, data privacy, computation, and domain knowledge integration.
- Suggest future directions and research pathways for MFMs in clinical practice.
Proposed method
- Define foundational models and summarize their pre-training objectives (contrastive, generative, hybrid).
- Classify medical FMs into six groups: VPM-Generalist, TPM-Hybrid, TPM-Contrastive, TPM-Generative, VPM-Adaptations, TPM-Conversational.
- Review representative works for textually prompted and visually prompted models, with emphasis on medical imaging tasks.
- Discuss prompts engineering, instruction-aligning, and how LM-inspired prompting translates to vision-language medical tasks.
- Aggregate open-source implementations and provide references for ongoing updates (GitHub initiative).

Experimental results
Research questions
- RQ1What are the fundamental concepts and training strategies enabling medical foundation models?
- RQ2How can medical imaging FMs be categorized by prompting and modality, and what are representative approaches per category?
- RQ3What are the key clinical opportunities, limitations, and challenges (interpretability, privacy, computation, distribution shifts) for MFMs?
- RQ4What future directions and open research questions will guide MFMs toward broader clinical impact?
Key findings
- FMs enable multi-modality interpretation and in-context adaptation with limited labeled data.
- A six-group taxonomy for medical FMs covers textually and visually prompted, including conversational and generalist variants.
- Prominent examples (e.g., MedCLIP, BiomedCLIP, MI-Zero, BioViL-T, Med-Flamingo) illustrate zero-shot, few-shot, and multimodal capabilities in radiology, pathology, and other domains.
- Prompting, instruction-alignment, and domain knowledge integration are central to improving medical imaging FM performance.
- Privacy-preserving data use and federated learning are highlighted as key advantages in medical settings.
- The survey emphasizes ongoing updates and open-source resources for rapid dissemination and evaluation.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.