Skip to main content
QUICK REVIEW

[Paper Review] eXplainable Artificial Intelligence on Medical Images: A Survey

Matteus Vargas Simão da Silva, Rodrigo Reis Arrais|arXiv (Cornell University)|May 12, 2023
Radiomics and Machine Learning in Medical Imaging4 citations
TL;DR

This survey provides a comprehensive analysis of explainable AI (XAI) techniques applied to medical imaging, evaluating 32 recent studies that use methods like LIME and SHAP to interpret deep learning models. It highlights the dominance of CNNs and the growing need for standardized evaluation metrics and shared medical image datasets to improve model transparency and clinical trustworthiness.

ABSTRACT

Over the last few years, the number of works about deep learning applied to the medical field has increased enormously. The necessity of a rigorous assessment of these models is required to explain these results to all people involved in medical exams. A recent field in the machine learning area is explainable artificial intelligence, also known as XAI, which targets to explain the results of such black box models to permit the desired assessment. This survey analyses several recent studies in the XAI field applied to medical diagnosis research, allowing some explainability of the machine learning results in several different diseases, such as cancers and COVID-19.

Motivation & Objective

  • To analyze recent advancements in explainable AI (XAI) for medical image diagnosis, focusing on interpretability and trust in deep learning models.
  • To identify common XAI techniques, model architectures, and evaluation practices across medical imaging applications.
  • To highlight critical challenges such as lack of standardization in XAI evaluation, data scarcity, and reliance on private or non-reproducible datasets.
  • To assess the role of XAI in improving clinical decision-making by providing interpretable, trustworthy, and actionable model explanations.
  • To advocate for the creation of a unified, large-scale, and diverse medical image database to support pretraining and model generalization.

Proposed method

  • Systematic review of 32 recent studies (2020–2023) in XAI applied to medical imaging, selected for novelty, clinical relevance, and use of explainable deep learning techniques.
  • Categorization of models by architecture (e.g., ResNet, DenseNet, Vision Transformers, U-Net), disease targets, and XAI frameworks (LIME, SHAP, saliency maps).
  • Qualitative and quantitative analysis of XAI evaluation practices, including human expert validation and model-agnostic vs. model-specific explanation methods.
  • Evaluation of dataset characteristics, including data availability (public vs. private), data size, and modality (e.g., X-ray, MRI, fundus, CT).
  • Identification of trends in XAI adoption, such as the widespread use of LIME and SHAP for local feature attribution in image classification.
  • Analysis of model performance and interpretability trade-offs, particularly in high-stakes conditions like pneumonia and COVID-19 detection.

Experimental results

Research questions

  • RQ1Which XAI techniques are most commonly used in medical image deep learning, and how do they support clinical interpretability?
  • RQ2What are the dominant deep learning architectures used in medical image XAI, and how do they compare in performance and explainability?
  • RQ3To what extent do current XAI methods in medical imaging support reliable, reproducible, and trustworthy model interpretation across diverse clinical scenarios?
  • RQ4What are the major challenges in evaluating XAI methods in medical imaging, and why is there no standardized metric for explainability assessment?
  • RQ5How do data availability, dataset privacy, and lack of large-scale shared databases affect the development and deployment of XAI in clinical settings?

Key findings

  • LIME and SHAP were the most widely used XAI frameworks, appearing in approximately 40% of the reviewed studies (7 for LIME, 6 for SHAP), primarily to explain the contribution of image regions to predictions.
  • Convolutional Neural Networks (CNNs) were the most prevalent model architecture, used in about 56% of the reviewed models, due to their effectiveness in medical image feature extraction.
  • COVID-19 and pneumonia were the most targeted pathologies, accounting for 25% of the studied conditions, reflecting the urgent need for explainable AI during the pandemic.
  • Despite high accuracy, many models remain clinically untrustworthy due to their 'black-box' nature, underscoring the need for interpretability to support clinical decision-making.
  • A significant number of datasets used in the studies were private, limiting reproducibility and hindering large-scale validation and benchmarking.
  • The lack of a standardized evaluation metric for XAI effectiveness was a recurring issue, with each study defining interpretability differently based on context and user needs.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.