[論文レビュー] Computation-Efficient Era: A Comprehensive Survey of State Space Models in Medical Image Analysis
本サーベイは、特にMambaアーキテクチャを含む状態空間モデル(SSMs)について、医用画像解析分野における包括的レビューを提供している。本研究では、画像モodalitiy、臓器、およびセグメンテーション、分類、マルチモodal融合などのタスクごとに、Mambaベースの手法を体系的に分類し、CNN やトランスフォーマーの長距離依存性モデリングやリソース制約下での限界を克服する可能性を強調している。計算効率の高さも顕著である。
Sequence modeling plays a vital role across various domains, with recurrent neural networks being historically the predominant method of performing these tasks. However, the emergence of transformers has altered this paradigm due to their superior performance. Built upon these advances, transformers have conjoined CNNs as two leading foundational models for learning visual representations. However, transformers are hindered by the $\mathcal{O}(N^2)$ complexity of their attention mechanisms, while CNNs lack global receptive fields and dynamic weight allocation. State Space Models (SSMs), specifically the extit{ extbf{Mamba}} model with selection mechanisms and hardware-aware architecture, have garnered immense interest lately in sequential modeling and visual representation learning, challenging the dominance of transformers by providing infinite context lengths and offering substantial efficiency maintaining linear complexity in the input sequence. Capitalizing on the advances in computer vision, medical imaging has heralded a new epoch with Mamba models. Intending to help researchers navigate the surge, this survey seeks to offer an encyclopedic review of Mamba models in medical imaging. Specifically, we start with a comprehensive theoretical review forming the basis of SSMs, including Mamba architecture and its alternatives for sequence modeling paradigms in this context. Next, we offer a structured classification of Mamba models in the medical field and introduce a diverse categorization scheme based on their application, imaging modalities, and targeted organs. Finally, we summarize key challenges, discuss different future research directions of the SSMs in the medical domain, and propose several directions to fulfill the demands of this field. In addition, we have compiled the studies discussed in this paper along with their open-source implementations on our GitHub repository.
研究の動機と目的
- 状態空間モデル(SSMs)、特にMambaを対象とした、医用画像解析分野における体系的かつ最新のレビューを提供すること。
- CNNが長距離依存性を捉えることの限界と、トランスフォーマーの計算複雑性の問題に対処すること。
- 応用分野、画像モダリティ、標的臓器に応じてMambaベースの手法を分類・分析すること。
- 説明可能性、分類タスクにおける性能ギャップ、医用画像分野における基礎モデルの欠如といった主な課題を特定すること。
- マルチモーダル学習、医用基礎モデル、2D/3Dデータ向けの改善済み走査スキームを含む、今後の研究方向性を提案すること。
提案手法
- 本サーベイは、Mambaの選択的状態空間メカニズムとハードウェアに配慮したアーキテクチャに焦点を当てた、SSMsの理論的レビューを実施している。
- 応用タイプ(例:セグメンテーション、分類)、画像モダリティ(MRI、CT、レントゲン)、標的臓器に基づいた、Mambaモデルの新規分類法を提唱している。
- Mambaのシーケンス長に対する線形複雑度(O(N))を、トランスフォーマーのO(N²)自己注意メカニズムと比較して分析している。
- 画像再構成、レジストレーション、マルチモーダル理解などの医用画像タスクにおいて、Mambaの性能を評価している。
- SSMとアテンションメカニズムを統合したMamba-2アーキテクチャなど、アーキテクチャ的イノベーションを検討している。
- 再現可能性とアクセス性を高めるために、オープンソース実装と関連論文を定期的に更新するGitHubリポジトリにまとめている。

実験結果
リサーチクエスチョン
- RQ1Mambaモデルは、計算効率を維持したまま、医用画像における長距離依存性モデリングにおいて、CNN やトランスフォーマーと比べてどのように異なるか?
- RQ2Mambaの性能を可能にしている主なアーキテクチャ的要素は何か?
- RQ3Mambaモデルが特に有望である医用画像応用分野(例:セグメンテーション、分類、再構成)は何か?
- RQ4臨床的AI分野におけるMambaの採用を妨げる主な課題(説明可能性、高レベルビジョンタスクにおける性能ギャップなど)は何か?
- RQ5マルチモーダル学習や基礎モデルの分野において、医用画像分野でのMambaベースのモデルを進歩させるために、今後最も有望な研究方向性は何か?
主な発見
- Mambaモデルは、シーケンス長に対して線形複雑度(O(N))を達成しており、トランスフォーマーのO(N²)自己注意メカニズムに比べて顕著な効率的利点を有する。
- 自己回帰的および長シーケンスタスクでは強力な性能を示すが、画像分類タスクでは現在のSOTA CNN やViTモデルに比べて性能が劣っている。
- SSMとアテンションを統合したMamba-2アーキテクチャは、理論的にSSMとアテンションメカニズムのギャップを埋める仕組みを有し、性能向上を実現している。
- 長距離的な空間的依存性が重要なタスク(例:画像セグメンテーション、再構成、レジストレーション)において、Mambaモデルは強く有望な性能を示している。
- 医用画像向けに特化したMambaベースの基礎モデルやマルチモーダル学習フレームワークの開発において、依然として大きな研究ギャップが存在する。
- Mambaの視覚的タスクにおける説明可能性は限定的であり、医療画像解析における意思決定プロセスへの理論的・実証的洞察は部分的である。

より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。