Skip to main content
QUICK REVIEW

[Paper Review] A Survey on Hallucination in Large Vision-Language Models

Hanchao Liu, Wenyuan Xue|arXiv (Cornell University)|Feb 1, 2024
Brain Tumor Detection and Classification20 citations
TL;DR

This survey defines LVLM hallucinations, catalogs their symptoms, reviews evaluation benchmarks and mitigation methods, and discusses causes and future directions.

ABSTRACT

Recent development of Large Vision-Language Models (LVLMs) has attracted growing attention within the AI landscape for its practical implementation potential. However, ``hallucination'', or more specifically, the misalignment between factual visual content and corresponding textual generation, poses a significant challenge of utilizing LVLMs. In this comprehensive survey, we dissect LVLM-related hallucinations in an attempt to establish an overview and facilitate future mitigation. Our scrutiny starts with a clarification of the concept of hallucinations in LVLMs, presenting a variety of hallucination symptoms and highlighting the unique challenges inherent in LVLM hallucinations. Subsequently, we outline the benchmarks and methodologies tailored specifically for evaluating hallucinations unique to LVLMs. Additionally, we delve into an investigation of the root causes of these hallucinations, encompassing insights from the training data and model components. We also critically review existing methods for mitigating hallucinations. The open questions and future directions pertaining to hallucinations within LVLMs are discussed to conclude this survey.

Motivation & Objective

  • Clarify the concept of hallucinations in LVLMs and categorize symptoms (judgment vs description) and semantic facets (object, attribute, relation).
  • Review LVLM-specific evaluation methods and benchmarks for non-hallucinatory generation and hallucination discrimination.
  • Analyze root causes from data, vision encoders, modality alignment, and LLM components to guide mitigation.
  • Survey existing mitigation approaches across data, vision, connection modules, decoding, and post-processing.
  • Discuss open questions and future directions to advance reliable LVLMs.

Proposed method

  • Define hallucination in LVLMs and present a taxonomy of symptoms (object, attribute, relation; judgment vs description).
  • Classify evaluation approaches into non-hallucinatory generation and hallucination discrimination, with discriminative vs generative benchmarks.
  • Summarize causes from data quality, vision encoder limitations, modality alignment, and LLM-induced factors.
  • Review mitigation strategies including data curation, vision encoder scaling, improved connection modules, decoding optimization, and post-processing.

Experimental results

Research questions

  • RQ1What constitutes hallucination in LVLMs and how can it be systematically categorized?
  • RQ2How are LVLM hallucinations evaluated, and what benchmarks exist for discrimination and generation tasks?
  • RQ3What are the primary data, model, and integration factors causing LVLM hallucinations?
  • RQ4What mitigation strategies are effective against LVLM hallucinations across data, architecture, and decoding?

Key findings

  • Hallucinations in LVLMs include object, attribute, and relation errors beyond simple object presence.
  • Evaluation methods for LVLM hallucinations split into non-hallucinatory generation and hallucination discrimination, with corresponding discriminative and generative benchmarks.
  • Benchmarks exist with varying sizes and metrics, highlighting a focus on object-level evaluation in discriminative tasks and broader hallucination categories in generative tasks.
  • Causes of LVLM hallucinations arise from data biases and annotation issues, vision encoder limitations, modality alignment gaps, and LLM-related factors such as context attention and decoding randomness.
  • Mitigation strategies span data optimization, vision encoder scaling and perceptual enhancements, improved connection modules, decoding-level adjustments, and post-processing methods.
  • There is a recognized need for richer supervision, multi-modality integration, LVLMs as agents, and interpretability-focused research to reduce hallucinations.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.