Skip to main content
QUICK REVIEW

[Paper Review] Explainability for Vision Foundation Models: A Survey

Rémi Kazmierczak, Eloïse Berthier|ArXiv.org|Jan 21, 2025
Organizational Management and Leadership3 citations
TL;DR

This survey maps the intersection of vision foundation models (PFMs) and explainable AI (XAI), compiling 122 papers, categorizing methods into inherently explainable and post-hoc approaches, and outlining evaluation practices and future challenges.

ABSTRACT

As artificial intelligence systems become increasingly integrated into daily life, the field of explainability has gained significant attention. This trend is particularly driven by the complexity of modern AI models and their decision-making processes. The advent of foundation models, characterized by their extensive generalization capabilities and emergent uses, has further complicated this landscape. Foundation models occupy an ambiguous position in the explainability domain: their complexity makes them inherently challenging to interpret, yet they are increasingly leveraged as tools to construct explainable models. In this survey, we explore the intersection of foundation models and eXplainable AI (XAI) in the vision domain. We begin by compiling a comprehensive corpus of papers that bridge these fields. Next, we categorize these works based on their architectural characteristics. We then discuss the challenges faced by current research in integrating XAI within foundation models. Furthermore, we review common evaluation methodologies for these combined approaches. Finally, we present key observations and insights from our survey, offering directions for future research in this rapidly evolving field.

Motivation & Objective

  • Summarize how vision foundation models are used to enable XAI in computer vision.
  • Catalog and classify XAI methods that leverage pretrained foundation models (PFMs) in vision tasks.
  • Evaluate how explanations are produced, stored, and assessed for PFMs in vision.
  • Identify gaps, challenges, and future research directions in integrating XAI with vision PFMs.

Proposed method

  • Assemble a corpus of 122 studies bridging XAI and PFMs in vision.
  • Provide taxonomy of XAI methods grounded in the PFMs context.
  • Differentiate two primary categories: inherently explainable models and post-hoc methods.
  • Discuss evaluation methodologies for explanations and their limitations.
  • Highlight challenges and open questions to guide future work.

Experimental results

Research questions

  • RQ1How are vision PFMs used to facilitate XAI methods (inherently explainable vs post-hoc)?
  • RQ2What are the prevailing evaluation methodologies for explanations produced with PFMs in vision?
  • RQ3What challenges and open problems exist at the intersection of PFMs and XAI in vision?
  • RQ4What future directions are suggested for advancing XAI with vision PFMs?

Key findings

  • A comprehensive corpus of 122 studies bridging XAI and foundation models in vision is compiled.
  • The taxonomy organizes methods into inherently explainable models and post-hoc explanations, with a detailed breakdown of subtypes.
  • PFMs enable XAI by embedding meaningful concepts and leveraging multimodal capabilities (e.g., CLIP, Grounding DINO) for explanations.
  • Evaluation trends include both qualitative and quantitative metrics, with notable reliance on dataset-specific benchmarks and meta-explanations.
  • The survey identifies core challenges and open questions, offering directions for future research in integrating PFMs with XAI in vision.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.