[논문 리뷰] Image Segmentation in Foundation Model Era: A Survey
기초 모델이 일반적이고 프롬프트 가능한 이미지 분할을 어떻게 가능하게 하는지에 대한 포괄적 고찰로, CLIP, 확산 모델, DINO, SAM의 발생하는 분할 지식을 강조하고, 남은 과제와 향후 방향을 제시합니다.
Image segmentation is a long-standing challenge in computer vision, studied continuously over several decades, as evidenced by seminal algorithms such as N-Cut, FCN, and MaskFormer. With the advent of foundation models (FMs), contemporary segmentation methodologies have embarked on a new epoch by either adapting FMs (e.g., CLIP, Stable Diffusion, DINO) for image segmentation or developing dedicated segmentation foundation models (e.g., SAM). These approaches not only deliver superior segmentation performance, but also herald newfound segmentation capabilities previously unseen in deep learning context. However, current research in image segmentation lacks a detailed analysis of distinct characteristics, challenges, and solutions associated with these advancements. This survey seeks to fill this gap by providing a thorough review of cutting-edge research centered around FM-driven image segmentation. We investigate two basic lines of research -- generic image segmentation (i.e., semantic segmentation, instance segmentation, panoptic segmentation), and promptable image segmentation (i.e., interactive segmentation, referring segmentation, few-shot segmentation) -- by delineating their respective task settings, background concepts, and key challenges. Furthermore, we provide insights into the emergence of segmentation knowledge from FMs like CLIP, Stable Diffusion, and DINO. An exhaustive overview of over 300 segmentation approaches is provided to encapsulate the breadth of current research efforts. Subsequently, we engage in a discussion of open issues and potential avenues for future research. We envisage that this fresh, comprehensive, and systematic survey catalyzes the evolution of advanced image segmentation systems. A public website is created to continuously track developments in this fast advancing field: \url{https://github.com/stanley-313/ImageSegFM-Survey}.
연구 동기 및 목표
- 기초 모델이 이미지 분할을 일반적이고 프롬프트 가능한 작업으로 어떻게 변환시키는지 설명한다.
- CLIP, 확산 모델, DINO, SAM 및 다른 FM들이 분할 지식의 발생에 어떻게 기여하는지 검토한다.
- 개방- 및 폐쇄-어휘 설정에서 GIS와 프롬프트 가능한 분할 접근법을 분류하고 분석한다.
- 다양한 작업에 걸친 다양한 프롬프트를 통일하기 위한 학습 없는 분할 및 프롬프트 전략을 논의한다.
- 향후 FM 기반 분할 연구를 이끌어 갈 개방 이슈와 가능한 방향을 식별한다.
제안 방법
- 입력에서 마스크와 어휘로의 매핑 f로서의 분할에 대한 통합된 수학적 공식화를 제공한다.
- 일반 이미지 분할과 프롬프트 가능한 이미지 분할을 구분하는 분류학을 개발한다.
- CLIP, 확산 모델, DINO, SAM으로부터의 FM과 지식 발생 기제를 조사한다.
- FM으로부터 분할 능력을 전이하거나 추출하는 방법을 철저히 검토한다.
- 제로샷 및 파샷 기능을 가능하게 하는 학습 없는 분할 및 프롬프트 기반 인터페이스를 논의한다.

실험 결과
연구 질문
- RQ1기초 모델이 GIS와 PIS 작업 전반에서 범용적이고 프롬프트 가능한 분할을 어떻게 가능하게 하는가?
- RQ2CLIP, DINO, 확산 모델, SAM으로부터의 분할 지식이 발생하도록 하는 메커니즘은 무엇인가?
- RQ3FM을 활용한 개방 어휘(Open-vocabulary) 및 학습-free 분할을 달성하기 위한 효과적인 전략은 무엇인가?
- RQ4FM 기반 이미지 분할의 주요 개방 이슈와 향후 방향은 무엇인가?
주요 결과
- 기초 모델은 여러 분할 작업에 대해 프롬프트 가능하고 다양한 분할 작업에 적응할 수 있는 분할 일반가를 가능하게 한다.
- 분할 지식은 CLIP, DINO, 확산 모델의 정렬(alignment) 및 어텐션 메커니즘과 SAM의 교차 어텐션에서 발생할 수 있다.
- 학습 없는 및 소수 샷 프롬프트 접근은 작업별 학습 없이 제로샷 또는 파샷 분할을 가능하게 한다.
- 개방 어휘 분할은 FM의 능력으로 인해 점차 강조되며, 폐쇄 어휘 설정을 넘어 확장되고 있다.
- 의미적, 인스턴스, 판노픽 작업에 걸친 FM 기반 분할 기법의 스펙트럼이 있으며, 실용적인 대화형 인터랙티브 및 텍스트 기반 프롬프트를 포함한다.
![Figure 2: MLLMs driven solutions lead to more powerful pixel reasoning and understanding capabilities, e.g . , multi-target reasoning segmentation, instance segmentation with text descriptions, referring segmentation and conversation. (Figure adapted courtesy of [ 60 ] )](https://ar5iv.labs.arxiv.org/html/2408.12957/assets/x2.png)
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.