[論文レビュー] Image Segmentation in Foundation Model Era: A Survey
基盤モデルが一般的でプロンプト可能な画像分割をどのように実現するかについての包括的な調査。CLIP、diffusion models、DINO、および SAM からの新たな分割知識の出現を強調し、未解決の課題と今後の方向性を概説する。
Image segmentation is a long-standing challenge in computer vision, studied continuously over several decades, as evidenced by seminal algorithms such as N-Cut, FCN, and MaskFormer. With the advent of foundation models (FMs), contemporary segmentation methodologies have embarked on a new epoch by either adapting FMs (e.g., CLIP, Stable Diffusion, DINO) for image segmentation or developing dedicated segmentation foundation models (e.g., SAM). These approaches not only deliver superior segmentation performance, but also herald newfound segmentation capabilities previously unseen in deep learning context. However, current research in image segmentation lacks a detailed analysis of distinct characteristics, challenges, and solutions associated with these advancements. This survey seeks to fill this gap by providing a thorough review of cutting-edge research centered around FM-driven image segmentation. We investigate two basic lines of research -- generic image segmentation (i.e., semantic segmentation, instance segmentation, panoptic segmentation), and promptable image segmentation (i.e., interactive segmentation, referring segmentation, few-shot segmentation) -- by delineating their respective task settings, background concepts, and key challenges. Furthermore, we provide insights into the emergence of segmentation knowledge from FMs like CLIP, Stable Diffusion, and DINO. An exhaustive overview of over 300 segmentation approaches is provided to encapsulate the breadth of current research efforts. Subsequently, we engage in a discussion of open issues and potential avenues for future research. We envisage that this fresh, comprehensive, and systematic survey catalyzes the evolution of advanced image segmentation systems. A public website is created to continuously track developments in this fast advancing field: \url{https://github.com/stanley-313/ImageSegFM-Survey}.
研究の動機と目的
- 基盤モデルが画像分割を汎用的でプロンプト可能なタスクに変換する方法を説明する。
- CLIP、diffusion models、DINO、SAM、およびその他のFMを分割知識の出現の観点からレビューする。
- GISとプロンプト可能な分割アプローチをオープン語彙設定とクローズド語彙設定の下で分類・分析する。
- トレーニング不要の分割とプロンプト戦略を用いて、タスク間で多様なプロンプトを統一する。
- FMベースの分割研究の将来の方向性と未解決の課題を特定する。)
提案手法
- 入力からマスクと語彙への写像 f としての分割の統一的な数理定式化を提供する。
- 一般的な画像分割とプロンプト可能な画像分割を区別する分類学を開発する。
- CLIP、diffusion models、DINO、SAM からのFMと知識出現メカニズムを調査する。
- FM から分割能力を転移・抽出する方法を網羅的にレビューする。
- トレーニング不要の分割とゼロショット・少数ショット機能を可能にするプロンプトベースのインターフェースについて論じる。

実験結果
リサーチクエスチョン
- RQ1基盤モデルはGISとPISタスクを横断する汎用でプロンプト可能な分割をどのように実現するか?
- RQ2CLIP、DINO、diffusion models、SAM から分割知識が出現する仕組みは何か?
- RQ3FMを用いてオープン語彙とトレーニング不要の分割を達成する効果的な戦略は何か?
- RQ4FMベースの画像分割における主要な未解決課題と今後の方向性は何か?
主な発見
- 基盤モデルは、複数の分割タスクに適応可能でプロンプト可能な分割のジェネラリストを実現する。
- 分割知識は、CLIP、DINO、diffusion models の整合性・アテンション機構や SAM のクロスアテンションから出現し得る。
- トレーニング不要および少数ショットプロンプティングのアプローチは、タスク固有の訓練なしでゼロショットまたは少数ショット分割を可能にする。
- オープン語彙分割はFMの能力のためますます重視され、クローズド語彙設定を超えて拡大している。
- 意味的、インスタンス、パノプティックタスクにわたるFMベースの分割技術のスペクトルがあり、実践的なインタラクティブおよびテキスト駆動のプロンプトを備える。
![Figure 2: MLLMs driven solutions lead to more powerful pixel reasoning and understanding capabilities, e.g . , multi-target reasoning segmentation, instance segmentation with text descriptions, referring segmentation and conversation. (Figure adapted courtesy of [ 60 ] )](https://ar5iv.labs.arxiv.org/html/2408.12957/assets/x2.png)
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。