[Paper Review] Algorithm-Agnostic Explainability for Unsupervised Clustering
The paper introduces two global/local, algorithm-agnostic explainability methods (G2PC and L2PC) to explain unsupervised clustering across multiple algorithms, demonstrated on synthetic data and high-dimensional fMRI connectivity data.
Supervised machine learning explainability has developed rapidly in recent years. However, clustering explainability has lagged behind. Here, we demonstrate the first adaptation of model-agnostic explainability methods to explain unsupervised clustering. We present two novel "algorithm-agnostic" explainability methods - global permutation percent change (G2PC) and local perturbation percent change (L2PC) - that identify feature importance globally to a clustering algorithm and locally to the clustering of individual samples. The methods are (1) easy to implement and (2) broadly applicable across clustering algorithms, which could make them highly impactful. We demonstrate the utility of the methods for explaining five popular clustering methods on low-dimensional synthetic datasets and on high-dimensional functional network connectivity data extracted from a resting-state functional magnetic resonance imaging dataset of 151 individuals with schizophrenia and 160 controls. Our results are consistent with existing literature while also shedding new light on how changes in brain connectivity may lead to schizophrenia symptoms. We further compare the explanations from our methods to an interpretable classifier and find them to be highly similar. Our proposed methods robustly explain multiple clustering algorithms and could facilitate new insights into many applications. We hope this study will greatly accelerate the development of the field of clustering explainability.
Motivation & Objective
- Advance clustering explainability by adapting model-agnostic explainability to unsupervised clustering.
- Provide globally and locally interpretable feature importance across diverse clustering algorithms.
- Demonstrate utility on low-dimensional synthetic data and high-dimensional brain connectivity data from fMRI.
Proposed method
- Develop global permutation percent change (G2PC) to quantify feature importance across a clustering algorithm.
- Develop local perturbation percent change (L2PC) to quantify feature importance for individual samples.
- Show applicability across multiple clustering algorithms (algorithm-agnostic).
- Validate explanations on synthetic datasets and high-dimensional resting-state fMRI connectivity data.
- Compare explanations to an interpretable classifier to assess similarity.
Experimental results
Research questions
- RQ1Can model-agnostic explainability techniques be adapted to explain unsupervised clustering?
- RQ2Do G2PC and L2PC provide stable and meaningful feature importance both globally and per-sample across different clustering algorithms?
- RQ3Do explanations align with interpretations from an interpretable classifier and with existing brain connectivity literature?
Key findings
- G2PC and L2PC successfully identify feature importance globally and locally for several clustering methods.
- The methods are easy to implement and broadly applicable across clustering algorithms.
- Explanations on fMRI connectivity data from schizophrenia and control groups align with existing literature and offer new insights into brain connectivity changes.
- Explanations from the proposed methods show high similarity to an interpretable classifier.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.