[论文解读] A Survey on Hallucination in Large Vision-Language Models
本综述定义 LVLM 的幻觉,汇集其症状,回顾评估基准和缓解方法,并讨论原因与未来方向。
Recent development of Large Vision-Language Models (LVLMs) has attracted growing attention within the AI landscape for its practical implementation potential. However, ``hallucination'', or more specifically, the misalignment between factual visual content and corresponding textual generation, poses a significant challenge of utilizing LVLMs. In this comprehensive survey, we dissect LVLM-related hallucinations in an attempt to establish an overview and facilitate future mitigation. Our scrutiny starts with a clarification of the concept of hallucinations in LVLMs, presenting a variety of hallucination symptoms and highlighting the unique challenges inherent in LVLM hallucinations. Subsequently, we outline the benchmarks and methodologies tailored specifically for evaluating hallucinations unique to LVLMs. Additionally, we delve into an investigation of the root causes of these hallucinations, encompassing insights from the training data and model components. We also critically review existing methods for mitigating hallucinations. The open questions and future directions pertaining to hallucinations within LVLMs are discussed to conclude this survey.
研究动机与目标
- 明确 LVLMs 中幻觉的概念,并对症状进行分类(判断性 vs 描述性)以及语义要素(对象、属性、关系)。
- 回顾面向 LVLM 的特有评估方法与基准,涵盖非幻觉生成与幻觉判别。
- 从数据、视觉编码器、模态对齐以及大语言模型组件分析根本原因,以指导缓解。
- 调查现有缓解方法,涵盖数据、视觉、连接模块、解码与后处理。
- 讨论未解问题与未来方向,以推动更可靠的 LVLM。
提出的方法
- 定义 LVLMs 中的幻觉并给出症状分类法(对象、属性、关系;判断性 vs 描述性)。
- 将评估方法分为非幻觉生成与幻觉判别,并区分判别基准与生成基准。
- 总结来自数据质量、视觉编码器局限、模态对齐以及由 LLM 引起的因素的原因。
- 综述缓解策略,包括数据筛选与整理、视觉编码器规模提升、改进连接模块、解码优化和后处理。
实验结果
研究问题
- RQ1LVLMs 中的幻觉由何构成?如何系统地分类?
- RQ2LVLM 幻觉如何评估?判别与生成任务存在哪些基准?
- RQ3引发 LVLM 幻觉的主要数据、模型与集成因素是什么?
- RQ4针对数据、架构和解码,在 LVLM 幻觉方面有哪些有效的缓解策略?
主要发现
- LVLM 幻觉包括对象、属性和关系的错误,超出简单的对象存在。
- 评估方法分为非幻觉生成与幻觉判别,并具有相应的判别基准与生成基准。
- 基准存在,规模与指标不一,突出在判别任务中关注对象级评估,在生成任务中覆盖更广的幻觉类别。
- 原因来自数据偏差和标注问题、视觉编码器局限、模态对齐差距,以及诸如上下文注意力和解码随机性等 LLM 相关因素。
- 缓解策略涵盖数据优化、视觉编码器规模与感知增强、改进连接模块、解码层面的调整和后处理方法。
- 已认识到需要更丰富的监督、多模态整合、将 LVLM 作为代理的应用,以及聚焦可解释性的研究以减少幻觉。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。