Skip to main content
QUICK REVIEW

[Paper Review] Are Disentangled Representations Helpful for Abstract Visual Reasoning?

Sjoerd van Steenkiste, Francesco Locatello|arXiv (Cornell University)|May 29, 2019
Explainable Artificial Intelligence (XAI)Computer Science79 references61 citations
TL;DR

The study shows that more disentangled representations learned unsupervisedly lead to better sample-efficient performance on abstract visual reasoning tasks, using 360 disentanglement models and 3600 reasoning models.

ABSTRACT

A disentangled representation encodes information about the salient factors of variation in the data independently. Although it is often argued that this representational format is useful in learning to solve many real-world down-stream tasks, there is little empirical evidence that supports this claim. In this paper, we conduct a large-scale study that investigates whether disentangled representations are more suitable for abstract reasoning tasks. Using two new tasks similar to Raven's Progressive Matrices, we evaluate the usefulness of the representations learned by 360 state-of-the-art unsupervised disentanglement models. Based on these representations, we train 3600 abstract reasoning models and observe that disentangled representations do in fact lead to better down-stream performance. In particular, they enable quicker learning using fewer samples.

Motivation & Objective

  • Motivate disentangled representations as a useful prior for downstream tasks beyond supervised settings.
  • Systematically evaluate a wide range of unsupervised disentanglement models on abstract reasoning tasks.
  • Assess whether disentanglement correlates with downstream performance across multiple metrics.
  • Quantify sample-efficiency gains as a function of disentanglement quality.

Proposed method

  • Construct two RPM-like abstract reasoning tasks based on dSprites and 3dshapes ground-truth factors.
  • Train 360 unsupervised disentanglement models spanning four approaches (β-VAE, FactorVAE, β-TCVAE, DIP-VAE).
  • Extract representations from these models and train 3600 Wild Relation Networks (WReN) to solve the reasoning tasks.
  • Measure downstream accuracy across training stages and correlate with five disentanglement metrics (BetaVAE, FactorVAE, MIG, DCI, SAP) and reconstruction error.
  • Compare few-shot versus many-shot regimes to assess sample efficiency.

Experimental results

Research questions

  • RQ1Do disentangled representations improve abstract visual reasoning performance compared to entangled representations?
  • RQ2How does downstream performance relate to different disentanglement metrics across sample regimes?
  • RQ3Is there a regime (few vs many samples) where disentanglement provides stronger benefits?

Key findings

  • More disentangled representations yield better sample-efficiency in the considered abstract visual reasoning tasks.
  • BetaVAE and FactorVAE scores show the strongest correlation with downstream accuracy in the few-sample regime; MIG and SAP correlate more weakly.
  • Reconstruction error correlates with downstream performance mainly in the many-sample regime and less so in the few-sample regime.
  • In the few-sample regime, disentanglement (e.g., FactorVAE score) remains strongly correlated with performance for both data sets (dSprites, 3dshapes).
  • In the many-sample regime, final accuracy shows less reliance on disentanglement metrics; reconstruction error becomes a stronger predictor of performance.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.