Skip to main content
QUICK REVIEW

[Paper Review] Generalized Out-of-Distribution Detection and Beyond in Vision Language Model Era: A Survey

Atsuyuki Miyai, Jingkang Yang|arXiv (Cornell University)|Jul 31, 2024
Text and Document Classification Technologies4 citations
TL;DR

This survey introduces Generalized Out-of-Distribution Detection v2, a unified framework for OOD detection, anomaly detection (AD), novelty detection (ND), open set recognition (OSR), and outlier detection (OD) in the Vision Language Model (VLM) and Large Vision Language Model (LVLM) era. It reveals that OOD detection and AD have become the dominant challenges due to paradigm shifts driven by models like CLIP and GPT-4V, and provides a comprehensive review of evolving definitions, benchmarks, methodologies, and future research directions in the VLM and LVLM era.

ABSTRACT

Detecting out-of-distribution (OOD) samples is crucial for ensuring the safety of machine learning systems and has shaped the field of OOD detection. Meanwhile, several other problems are closely related to OOD detection, including anomaly detection (AD), novelty detection (ND), open set recognition (OSR), and outlier detection (OD). To unify these problems, a generalized OOD detection framework was proposed, taxonomically categorizing these five problems. However, Vision Language Models (VLMs) such as CLIP have significantly changed the paradigm and blurred the boundaries between these fields, again confusing researchers. In this survey, we first present a generalized OOD detection v2, encapsulating the evolution of these fields in the VLM era. Our framework reveals that, with some field inactivity and integration, the demanding challenges have become OOD detection and AD. Then, we highlight the significant shift in the definition, problem settings, and benchmarks; we thus feature a comprehensive review of the methodology for OOD detection and related tasks to clarify their relationship to OOD detection. Finally, we explore the advancements in the emerging Large Vision Language Model (LVLM) era, such as GPT-4V. We conclude with open challenges and future directions. The resource is available at https://github.com/AtsuMiyai/Awesome-OOD-VLM.

Motivation & Objective

  • To unify and reframe the five closely related tasks—out-of-distribution (OOD) detection, anomaly detection (AD), novelty detection (ND), open set recognition (OSR), and outlier detection (OD)—within a single, evolving framework in the context of Vision Language Models (VLMs).
  • To analyze the paradigm shift caused by VLMs like CLIP and LVLMs like GPT-4V, which have blurred traditional boundaries between these tasks and redefined the core challenges in the field.
  • To clarify evolving definitions, problem settings, and benchmarks for OOD detection and related tasks, especially in light of VLM-based methods.
  • To conduct a systematic review of methodologies for OOD detection and related tasks in the VLM era, highlighting key techniques, baselines, and open problems.
  • To identify and outline future research directions, including unsolvable problem detection (UPD), real-world benchmarking, and theoretical understanding of LVLM behavior.

Proposed method

  • Proposes a new unified framework, Generalized OOD Detection v2, to integrate and reclassify OOD detection, AD, ND, OSR, and OD under a single taxonomy, reflecting the evolving landscape in the VLM era.
  • Reviews the evolution of OOD detection and related tasks using VLMs such as CLIP, analyzing shifts in definitions, problem settings, and benchmarking practices.
  • Categorizes and compares methodologies for OOD detection in the VLM era, including feature-based, logit-based, and uncertainty estimation techniques, with emphasis on CLIP-based and LVLM-based approaches.
  • Highlights the use of large pre-trained models and parameter-efficient fine-tuning (e.g., DSGF) for single-modal OOD detection, and advocates for broader adoption of pre-training in OOD research.
  • Introduces and evaluates real-world benchmarks such as ImageNet-ES and WILDS to bridge the gap between standard benchmarks and real-world data shifts.
  • Proposes future directions including unsolvable problem detection (UPD) using LVLM response perplexity, integration of UPD into diverse benchmarks like MuirBench, and theoretical analysis of LVLM robustness.

Experimental results

Research questions

  • RQ1How have the definitions and problem settings of OOD detection, AD, ND, OSR, and OD evolved in the VLM era, particularly with models like CLIP and GPT-4V?
  • RQ2What are the key methodological advancements in OOD detection using VLMs and LVLMs, and how do they differ from traditional single-modal approaches?
  • RQ3Why have some fields become inactive or integrated in the VLM era, and what are the dominant challenges now?
  • RQ4How can real-world benchmarks like ImageNet-ES and WILDS improve the evaluation of OOD detection in safety-critical applications?
  • RQ5What are the emerging research directions in the LVLM era, particularly regarding unsolvable problem detection (UPD), and how can they be formalized and evaluated?

Key findings

  • The generalized OOD detection v2 framework reveals that OOD detection and anomaly detection (AD) have emerged as the primary challenges in the VLM era, while other fields like ND and OSR have seen reduced distinctiveness or integration.
  • The advent of VLMs such as CLIP has significantly blurred the boundaries between OOD detection, AD, ND, OSR, and OD, necessitating a unified framework to clarify their relationships and distinctions.
  • CLIP-based OOD detection has become a dominant paradigm, with methods leveraging contrastive learning and zero-shot generalization, but real-world benchmarks like ImageNet-ES are needed to close the gap between standard benchmarks and real-world data shifts.
  • Parameter-efficient fine-tuning techniques like DSGF show promise for single-modal OOD detection when leveraging large pre-trained models, though this area remains underexplored compared to closed-set classification.
  • The emergence of LVLMs like GPT-4V and LLaVA has enabled new frontiers in OOD detection, including object-level detection and segmentation, especially when combined with methods like MCM.
  • Unsolvable problem detection (UPD) is identified as a critical emerging challenge, with potential solutions using response perplexity and model-agnostic post-hoc methods, and integration into diverse benchmarks like MuirBench is essential for robust evaluation.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.