Skip to main content
QUICK REVIEW

[论文解读] Navigating the landscape of multimodal AI in medicine: a scoping review on technical challenges and clinical applications

Daan Schouten, Giulia Nicoletti|arXiv (Cornell University)|Nov 6, 2024
Artificial Intelligence in Healthcare and Education被引用 5
一句话总结

对432项基于深度学习的多模态AI研究(2018–2024)在医学领域的范围性综述,详细介绍模态、架构、融合策略、挑战以及临床应用前景。

ABSTRACT

Recent technological advances in healthcare have led to unprecedented growth in patient data quantity and diversity. While artificial intelligence (AI) models have shown promising results in analyzing individual data modalities, there is increasing recognition that models integrating multiple complementary data sources, so-called multimodal AI, could enhance clinical decision-making. This scoping review examines the landscape of deep learning-based multimodal AI applications across the medical domain, analyzing 432 papers published between 2018 and 2024. We provide an extensive overview of multimodal AI development across different medical disciplines, examining various architectural approaches, fusion strategies, and common application areas. Our analysis reveals that multimodal AI models consistently outperform their unimodal counterparts, with an average improvement of 6.2 percentage points in AUC. However, several challenges persist, including cross-departmental coordination, heterogeneous data characteristics, and incomplete datasets. We critically assess the technical and practical challenges in developing multimodal AI systems and discuss potential strategies for their clinical implementation, including a brief overview of commercially available multimodal AI models for clinical decision-making. Additionally, we identify key factors driving multimodal AI development and propose recommendations to accelerate the field's maturation. This review provides researchers and clinicians with a thorough understanding of the current state, challenges, and future directions of multimodal AI in medicine.

研究动机与目标

  • 对2018–2024年跨学科与任务的医学领域基于深度学习的多模态AI现状进行调查。
  • 描述在多模态医疗AI中使用的数据模态、体系结构和融合策略。
  • 识别包括数据可用性、缺失模态以及验证实践在内的技术和实际挑战。
  • 讨论临床实施的路径,包括监管、可解释性和数据获取等考虑因素。
  • 提供加速医疗保健领域多模态AI成熟的建议。

提出的方法

  • 对2018年至2024年发表的432篇论文进行系统范围界定综述。
  • 纳入标准:深度神经网络、来自不同医疗专科的多模态数据,以及特定医疗任务。
  • 数据源分析包括模态分类与器官-系统映射;对公开数据集与私有数据集的评估。
  • 对报告的性能提升和验证实践进行定量综合。
  • 对融合策略、编码器架构以及缺失模态处理进行批判性评估。
Figure 1: Overview of the screening process.
Figure 1: Overview of the screening process.

实验结果

研究问题

  • RQ1在医学多模态AI研究中,常见的数据模态及模态组合有哪些?
  • RQ2哪些器官系统和医疗任务在多模态AI研究中占主导地位,与单模态基线相比的典型性能提升是多少?
  • RQ3最常见的架构选择与融合策略有哪些,缺失数据如何处理?
  • RQ4临床采用和数据共享的主要障碍是什么,如何应对?
  • RQ5推动多模态AI发展的因素有哪些,哪些建议可以加速其成熟?

主要发现

  • 432项研究(2018–2024)显示多模态模型在子集分析中优于单模态对应物,平均AUC提高6.2个百分点。
  • 大多数研究(82%)使用内部验证;只有少数采用外部验证。
  • 放射学和文本模态最常见(各约30%),放射学/文本是最频繁的组合(206个实例)。
  • CNN在编码器中占主导地位(82%),中间融合是最常见的融合阶段(79%),连接(拼接)仍是主流融合方法(69%)。
  • 早期融合很罕见(6%),而晚期融合(14%)通常利用单模态预测或为每种模态设置的独立模型;基于注意力的中间融合使用日益增加。
  • 公共数据集被大量使用(61%),但私有数据集(24%)和有限的外部验证是显著的空白点。
  • 处理缺失模态是一个主要挑战;69%的论文排除了不完整的条目,而基于学习的插补和灵活的架构被作为替代方案进行探索。
Figure 2: Overview of the data modalities used in the reviewed articles. (A) Distribution of articles by year. Bar chart shows an exponential increment in the number of studies per year from 2018 to 2024. Extrapolating, the number of multimodal medical AI studies is expected to reach 199 by the end
Figure 2: Overview of the data modalities used in the reviewed articles. (A) Distribution of articles by year. Bar chart shows an exponential increment in the number of studies per year from 2018 to 2024. Extrapolating, the number of multimodal medical AI studies is expected to reach 199 by the end

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。