Skip to main content
QUICK REVIEW

[Paper Review] Disentangling Factors of Variation Using Few Labels

Francesco Locatello, Michael Tschannen|arXiv (Cornell University)|May 3, 2019
Spectroscopy and Chemometric Analyses70 references50 citations
TL;DR

The paper shows that a very small amount of labeled data can guide both unsupervised and semi-supervised disentanglement learning, enabling reliable disentangled representations and effective model selection across large-scale experiments (over 52,000 models).

ABSTRACT

Learning disentangled representations is considered a cornerstone problem in representation learning. Recently, Locatello et al. (2019) demonstrated that unsupervised disentanglement learning without inductive biases is theoretically impossible and that existing inductive biases and unsupervised methods do not allow to consistently learn disentangled representations. However, in many practical settings, one might have access to a limited amount of supervision, for example through manual labeling of (some) factors of variation in a few training examples. In this paper, we investigate the impact of such supervision on state-of-the-art disentanglement methods and perform a large scale study, training over 52000 models under well-defined and reproducible experimental conditions. We observe that a small number of labeled examples (0.01--0.5\% of the data set), with potentially imprecise and incomplete labels, is sufficient to perform model selection on state-of-the-art unsupervised models. Further, we investigate the benefit of incorporating supervision into the training process. Overall, we empirically validate that with little and imprecise supervision it is possible to reliably learn disentangled representations.

Motivation & Objective

  • Motivate practical disentanglement learning with limited supervision.
  • Quantify how few labels affect model selection and training for state-of-the-art methods.
  • Assess robustness of supervision under label imprecision and partial labeling.
  • Provide practical guidelines for using limited supervision in disentangled representation learning.

Proposed method

  • Evaluate whether standard disentanglement metrics can identify good models using very few labels.
  • Train over 52,000 models across four datasets with 100 or 1000 labels under various labeling conditions.
  • Compare unsupervised training with supervised validation (U/S) to semi-supervised training with supervision during training (S2/S).
  • Incorporate a simple supervised regularizer into the loss to fuse labeled information into training (R_s) and assess its impact.
  • Use model selection metrics (MIG, DCI Disentanglement, SAP) and test metrics to evaluate disentanglement.
  • Examine robustness to imperfect labels (binned, noisy, partial) and label permutation.

Experimental results

Research questions

  • RQ1Can few labeled examples be sufficient to select good disentangled models from unsupervised training?
  • RQ2Does incorporating limited supervision into training outperform unsupervised training with supervised validation?
  • RQ3How robust are such supervision approaches to label noise, imprecision, and partial labeling?
  • RQ4Do results generalize across multiple standard disentanglement datasets?

Key findings

  • A small number of labels (0.01–0.5% of data) suffices to perform model selection for unsupervised disentanglement methods.
  • Unsupervised training with supervised validation enables reliable learning of disentangled representations.
  • Incorporating supervision during training often outperforms unsupervised training with validation alone.
  • Semi-supervised training is robust to label noise and partial/coarse labels.
  • Labeling more factors in a coarse way tends to help more than fine-grained labeling of few factors.
  • The approach provides practical guidelines for leveraging disentangled representations in real-world tasks.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.