[論文レビュー] Navigating the landscape of multimodal AI in medicine: a scoping review on technical challenges and clinical applications
医療分野における深層学習ベースのマルチモーダルAI研究432件のスコーピングレビュー(2018–2024)、モダリティ、アーキテクチャ、フュージョン戦略、課題、臨床導入の見通しを詳述。
Recent technological advances in healthcare have led to unprecedented growth in patient data quantity and diversity. While artificial intelligence (AI) models have shown promising results in analyzing individual data modalities, there is increasing recognition that models integrating multiple complementary data sources, so-called multimodal AI, could enhance clinical decision-making. This scoping review examines the landscape of deep learning-based multimodal AI applications across the medical domain, analyzing 432 papers published between 2018 and 2024. We provide an extensive overview of multimodal AI development across different medical disciplines, examining various architectural approaches, fusion strategies, and common application areas. Our analysis reveals that multimodal AI models consistently outperform their unimodal counterparts, with an average improvement of 6.2 percentage points in AUC. However, several challenges persist, including cross-departmental coordination, heterogeneous data characteristics, and incomplete datasets. We critically assess the technical and practical challenges in developing multimodal AI systems and discuss potential strategies for their clinical implementation, including a brief overview of commercially available multimodal AI models for clinical decision-making. Additionally, we identify key factors driving multimodal AI development and propose recommendations to accelerate the field's maturation. This review provides researchers and clinicians with a thorough understanding of the current state, challenges, and future directions of multimodal AI in medicine.
研究の動機と目的
- 2018–2024年の医療分野における深層学習ベースのマルチモーダルAIの景観を分野とタスク横断で調査する。
- マルチモーダル医療AIで使用されるデータモダリティ、アーキテクチャ的アプローチ、フュージョン戦略を特徴づける。
- データ取得可用性、欠損モダリティ、検証実践を含む技術的・実践的課題を特定する。
- 規制、説明可能性、データアクセスの考慮を含む臨床導入への道筋を論じる。
- 医療におけるマルチモーダルAIの成熟を加速するための推奨事項を提供する。
提案手法
- 2018年から2024年に公表された432件の論文のシステマティック・スコーピングレビュー。
- 収集基準:深層ニューラルネットワーク、異なる医療専門分野からのマルチモーダルデータ、および特定の医療タスク。
- データソース分析にはモダリティの分類と臓器系マッピングを含み、公開データセットと民間データセットの評価を含む。
- 報告された性能向上と検証実践の定量的統合。
- フュージョン戦略、エンコーダーアーキテクチャ、欠損モダリティの扱いの批判的評価。

実験結果
リサーチクエスチョン
- RQ1医療のマルチモーダルAI研究で用いられる一般的なデータモダリティとモダリティの組み合わせは何か。
- RQ2どの臓器系と医療タスクがマルチモーダルAI研究を支配し、単一モーダル基準と比較してどの程度の性能向上が一般的か。
- RQ3最も一般的なアーキテクチャの選択肢とフュージョン戦略は何で、欠損データはどう扱われているか。
- RQ4臨床導入とデータ共有の主な障害は何で、これをどう解決できるか。
- RQ5マルチモーダルAI開発を左右する要因は何で、それを成熟させる推奨は何か。
主な発見
- 432件(2018–2024)で、マルチモーダルモデルは単一モーダルの対照よりも優れており、部分分析で平均AUCが6.2ポイント改善。
- ほとんどの研究は内部検証を用いる(82%)、外部検証を採用するのは少数。
- 放射線診断とテキストモダリティが最も一般的(いずれも約30%)、放射線/テキストの組み合わせが最頻(206件)。
- エンコーダはCNNが支配的(82%)、中間結合が最も一般的(79%)、結合は依然として主要なフュージョン手法(69%)。
- 早結合は希少(6%)、後結合は(14%)が多く、単一モダリ値予測または各モダリティ用の別モデルを利用することが多い;注意機構を用いた中間結合が増加。
- 公開データセットが多用(61%)、ただし民間データセット(24%)と外部検証の不足が顕著。
- 欠損モダリティの扱いは大きな課題;論文の69%が不完全なエントリを除外、一方で学習ベースの補完と柔軟なアーキテクチャが代替として検討されている。

より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。