[Paper Review] Navigating the landscape of multimodal AI in medicine: a scoping review on technical challenges and clinical applications
A scoping review of 432 deep learning–based multimodal AI studies (2018–2024) across medicine, detailing modalities, architectures, fusion strategies, challenges, and clinical adoption prospects.
Recent technological advances in healthcare have led to unprecedented growth in patient data quantity and diversity. While artificial intelligence (AI) models have shown promising results in analyzing individual data modalities, there is increasing recognition that models integrating multiple complementary data sources, so-called multimodal AI, could enhance clinical decision-making. This scoping review examines the landscape of deep learning-based multimodal AI applications across the medical domain, analyzing 432 papers published between 2018 and 2024. We provide an extensive overview of multimodal AI development across different medical disciplines, examining various architectural approaches, fusion strategies, and common application areas. Our analysis reveals that multimodal AI models consistently outperform their unimodal counterparts, with an average improvement of 6.2 percentage points in AUC. However, several challenges persist, including cross-departmental coordination, heterogeneous data characteristics, and incomplete datasets. We critically assess the technical and practical challenges in developing multimodal AI systems and discuss potential strategies for their clinical implementation, including a brief overview of commercially available multimodal AI models for clinical decision-making. Additionally, we identify key factors driving multimodal AI development and propose recommendations to accelerate the field's maturation. This review provides researchers and clinicians with a thorough understanding of the current state, challenges, and future directions of multimodal AI in medicine.
Motivation & Objective
- Survey the landscape of deep learning–based multimodal AI in medicine across disciplines and tasks from 2018–2024.
- Characterize data modalities, architectural approaches, and fusion strategies used in multimodal medical AI.
- Identify technical and practical challenges, including data availability, missing modalities, and validation practices.
- Discuss pathways to clinical implementation, including regulatory, explainability, and data-access considerations.
- Provide recommendations to accelerate maturation of multimodal AI in healthcare.
Proposed method
- Systematic scoping review of 432 papers published between 2018 and 2024.
- Inclusion criteria: deep neural networks, multimodal data from different medical specialties, and specific medical tasks.
- Data source analysis includes modality categorization and organ-system mapping; evaluation of public vs private datasets.
- Quantitative synthesis of reported performance gains and validation practices.
- Critical assessment of fusion strategies, encoder architectures, and handling of missing modalities.

Experimental results
Research questions
- RQ1What are the prevalent data modalities and modality combinations used in medical multimodal AI studies?
- RQ2Which organ systems and medical tasks dominate multimodal AI research, and what are typical performance gains over unimodal baselines?
- RQ3What architectural choices and fusion strategies are most common, and how is missing data handled?
- RQ4What are the main obstacles to clinical adoption and data sharing, and how can these be addressed?
- RQ5What factors drive multimodal AI development, and what recommendations can accelerate its maturation?
Key findings
- 432 studies (2018–2024) show multimodal models outperform unimodal counterparts with an average AUC improvement of 6.2 percentage points in a subset analysis.
- Most studies (82%) use internal validation; only a minority employ external validation.
- Radiology and text modalities are the most common (each ~30%), with radiology/text being the most frequent combination (206 instances).
- CNNs dominate encoders (82%), with intermediate fusion being the most common fusion stage (79%), and concatenation remaining the prevalent fusion method (69%).
- Early fusion is rare (6%), while late fusion (14%) often leverages unimodal predictions or separate models for each modality; attention-based intermediate fusion is increasingly used.
- Public datasets are heavily utilized (61%), but private datasets (24%) and limited external validation are notable gaps.
- Handling missing modalities is a major challenge; 69% of papers exclude incomplete entries, while learning-based imputation and flexible architectures are explored as alternatives.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.