Skip to main content
QUICK REVIEW

[Paper Review] A Comprehensive Study on Visual Explanations for Spatio-temporal Networks.

Zhenqiang Li, Weimin Wang|arXiv (Cornell University)|May 1, 2020
Multimodal Machine Learning Applications34 references4 citations
TL;DR

This paper presents a comprehensive evaluation of gradient-based and perturbation-based attribution methods for spatio-temporal neural networks processing video inputs. It extends a 2D perturbation method to 3D video data, introduces complementary objective metrics, and reveals that attribution methods exhibit opposing performance on objective versus subjective evaluation criteria.

ABSTRACT

Identifying and visualizing regions that are significant for a given deep neural network model, i.e., attribution methods, is still a vital but challenging task, especially for spatio-temporal networks that process videos as input. Albeit some methods that have been proposed for video attribution, it is yet to be studied what types of network structures each video attribution method is suitable for. In this paper, we provide a comprehensive study of the existing video attribution methods of two categories, gradient-based and perturbation-based, for visual explanation of neural networks that take videos as the input (spatio-temporal networks). To perform this study, we extended a perturbation-based attribution method from 2D (images) to 3D (videos) and validated its effectiveness by mathematical analysis and experiments. For a more comprehensive analysis of existing video attribution methods, we introduce objective metrics that are complementary to existing subjective ones. Our experimental results indicate that attribution methods tend to show opposite performances on objective and subjective metrics.

Motivation & Objective

  • To investigate the suitability of existing video attribution methods—gradient-based and perturbation-based—for different spatio-temporal network architectures.
  • To extend a 2D perturbation-based attribution method to 3D video data through mathematical and empirical validation.
  • To introduce objective evaluation metrics that complement existing subjective assessments in video attribution.
  • To analyze the discrepancy between objective and subjective performance of attribution methods in video explanation tasks.

Proposed method

  • Extended a 2D perturbation-based attribution method to 3D by adapting the saliency computation to video tensors (spatio-temporal volumes).
  • Formalized the 3D perturbation method using a masking strategy over spatial and temporal dimensions to identify influential video segments.
  • Proposed new objective metrics based on model robustness and feature sensitivity to quantify attribution quality independently of human judgment.
  • Conducted controlled experiments across multiple video classification models to compare attribution methods under both objective and subjective evaluation.
  • Used mathematical analysis to validate the theoretical consistency of the extended 3D perturbation method.
  • Combined qualitative visualizations with quantitative metrics to assess attribution fidelity and stability.

Experimental results

Research questions

  • RQ1Which types of spatio-temporal network architectures are most suitable for gradient-based versus perturbation-based video attribution methods?
  • RQ2How can a 2D perturbation-based attribution method be effectively extended to 3D video data while preserving interpretability and accuracy?
  • RQ3To what extent do objective metrics correlate with subjective human evaluations of attribution quality in video explanation?
  • RQ4Do attribution methods show consistent performance across both objective and subjective evaluation criteria?

Key findings

  • The extended 3D perturbation-based method demonstrated consistent and stable attribution maps across diverse video datasets, validated through mathematical analysis and empirical results.
  • Objective metrics revealed that some attribution methods performed poorly despite receiving high subjective scores from human annotators.
  • A significant performance divergence was observed: methods scoring well on subjective evaluation often underperformed on objective metrics, and vice versa.
  • The study confirmed that relying solely on subjective evaluation may lead to misleading conclusions about attribution method quality.
  • The proposed objective metrics effectively captured model sensitivity and robustness, offering a reliable alternative to human judgment in attribution evaluation.
  • Gradient-based methods showed higher sensitivity to input perturbations compared to perturbation-based methods, especially in long-range temporal reasoning tasks.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.