[论文解读] A Survey of Hallucination in Large Foundation Models
对大型 foundation models(LFMs)在文本、图像、视频和音频中的幻觉现象进行全面综述,详细介绍类型、评估、检测、缓解、数据集和未来方向。
Hallucination in a foundation model (FM) refers to the generation of content that strays from factual reality or includes fabricated information. This survey paper provides an extensive overview of recent efforts that aim to identify, elucidate, and tackle the problem of hallucination, with a particular focus on ``Large'' Foundation Models (LFMs). The paper classifies various types of hallucination phenomena that are specific to LFMs and establishes evaluation criteria for assessing the extent of hallucination. It also examines existing strategies for mitigating hallucination in LFMs and discusses potential directions for future research in this area. Essentially, the paper offers a comprehensive examination of the challenges and solutions related to hallucination in LFMs.
研究动机与目标
- 对 LFMs 幻觉相关的现有工作进行分类。
- 在文本、图像、视频和音频模态下检视 LFMs。
- 总结检测方法、缓解策略、任务、数据集和评估指标。
- 提出未来研究方向和开源资源。
提出的方法
- 将 LFMs 分类为文本、图像、视频和音频四种模态。
- 评审幻觉的检测、缓解、任务、数据集和评估指标。
- 总结特定模态的研究与数据集(例如 HaluEval、Med-HALT、 M-HalDetect)。
- 强调用于缓解的提示、外部知识、知识对齐(grounding)和数据增强方法。
- 提供未来方向,包括自动化评估和经过筛选的知识源等未来方向。

实验结果
研究问题
- RQ1在不同模态下,LFMs 观察到的主要幻觉类型有哪些?
- RQ2为减少幻觉,提出了哪些 LFMs 的检测与缓解策略?
- RQ3在文本、图像、视频和音频中,用于评估幻觉的数据集和评估指标有哪些?
- RQ4哪些未来方向可以推动 LFMs 的可靠幻觉检测与缓解?
- RQ5领域特定和多语言的 LFMs 在幻觉特征与缓解方面有何差异?
主要发现
- 本综述提供了幻觉类型的分类法,以及 LFMs 在文本、图像、视频、音频等跨模态的概览。
- 它对各模态的检测方法、缓解技术、数据集和评估指标进行了编目。
- 它突出显示了诸如 HaluEval、Med-HALT、HALO、M-HalDetect 和 Lp-MusicCaps 等著名基准和数据集。
- 它讨论了基于 prompting、知识对齐和外部知识整合等缓解方向。
- 它强调了在自动化评估、知识整理和伦理考量方面的未来研究方向。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。