[论文解读] Explainability for Vision Foundation Models: A Survey
本综述映射视觉基础模型(PFMs)与可解释性AI(XAI)的交叉领域,汇编了122篇论文,将方法分为本质可解释与事后解释两大类,并概述评估实践与未来挑战。
As artificial intelligence systems become increasingly integrated into daily life, the field of explainability has gained significant attention. This trend is particularly driven by the complexity of modern AI models and their decision-making processes. The advent of foundation models, characterized by their extensive generalization capabilities and emergent uses, has further complicated this landscape. Foundation models occupy an ambiguous position in the explainability domain: their complexity makes them inherently challenging to interpret, yet they are increasingly leveraged as tools to construct explainable models. In this survey, we explore the intersection of foundation models and eXplainable AI (XAI) in the vision domain. We begin by compiling a comprehensive corpus of papers that bridge these fields. Next, we categorize these works based on their architectural characteristics. We then discuss the challenges faced by current research in integrating XAI within foundation models. Furthermore, we review common evaluation methodologies for these combined approaches. Finally, we present key observations and insights from our survey, offering directions for future research in this rapidly evolving field.
研究动机与目标
- 总结视觉基础模型如何用于在计算机视觉中实现XAI。
- 编目并分类在视觉任务中利用预训练基础模型(PFMs)的XAI方法。
- 评估PFMs在视觉领域的解释是如何产生、存储与评估的。
- 识别将XAI与视觉PFMs整合中的差距、挑战及未来研究方向。
提出的方法
- 汇集一个桥接XAI与视觉PFMs的122项研究语料库。
- 提供基于PFMs情境的XAI方法分类。
- 区分两大核心类别:本质可解释模型与事后解释。
- 讨论解释的评估方法及其局限性。
- 强调挑战与未解决的问题,为未来工作指明方向。
实验结果
研究问题
- RQ1视觉PFMs如何用于促进XAI方法(本质可解释与事后解释)?
- RQ2在视觉领域使用PFMs生成的解释的评估方法有哪些?
- RQ3PFMs与XAI在视觉领域交叉处存在哪些挑战和待解决的问题?
- RQ4为推进视觉PFMs的XAI发展,未来方向有哪些建议?
主要发现
- 编制出一个涵盖XAI与视觉基础模型的122项研究的综合语料库。
- 该分类将方法分为本质可解释模型与事后解释,并对子类型做了详细分解。
- PFMs通过嵌入有意义的概念并利用多模态能力(如CLIP、Grounding DINO)来实现解释。
- 评估趋势包括定性与定量指标,显著依赖于特定数据集的基准与元解释。
- 综述指出核心挑战与未解问题,并为将PFMs与XAI在视觉中的整合提供未来研究方向。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。