[Paper Review] Weakly-Supervised Disentanglement Without Compromises
The paper proves identifiability and provides adaptive weakly-supervised VAEs to learn disentangled representations from pairs of images lacking group annotations, and demonstrates strong downstream usefulness across tasks.
Intelligent agents should be able to learn useful representations by observing changes in their environment. We model such observations as pairs of non-i.i.d. images sharing at least one of the underlying factors of variation. First, we theoretically show that only knowing how many factors have changed, but not which ones, is sufficient to learn disentangled representations. Second, we provide practical algorithms that learn disentangled representations from pairs of images without requiring annotation of groups, individual factors, or the number of factors that have changed. Third, we perform a large-scale empirical study and show that such pairs of observations are sufficient to reliably learn disentangled representations on several benchmark data sets. Finally, we evaluate our learned representations and find that they are simultaneously useful on a diverse suite of tasks, including generalization under covariate shifts, fairness, and abstract reasoning. Overall, our results demonstrate that weak supervision enables learning of useful disentangled representations in realistic scenarios.
Motivation & Objective
- Motivate learning disentangled representations from non-i.i.d. image pairs where only a subset of factors change.
- Show that identifiability is achievable under weak assumptions with paired observations.
- Develop practical adaptive algorithms that infer factor sharing without group annotations.
- Demonstrate through large-scale experiments that weak supervision yields reliable disentanglement and useful representations.
Proposed method
- Propose a weakly-supervised generative model where two observations share a random subset of latent factors and differ in others.
- Derive identifiability results showing that constrained distribution matching yields disentangled posteriors up to permutation of latent axes.
- Introduce Ada-GVAE and Ada-ML-VAE variants that adaptively infer shared factors (S) from pairs by averaging posteriors of shared coordinates.
- Use a beta-VAE objective with a novel averaging constraint to enforce shared/non-shared coordinate consistency without known group structure.
- Estimate the number of shared factors k per pair with an elbow-method-like heuristic based on KL divergences between pair posteriors.
Experimental results
Research questions
- RQ1Can disentangled representations be identified from non-i.i.d. image pairs without explicit group annotations?
- RQ2How can we adaptively infer which latent factors are shared across a pair when k is unknown and variable?
- RQ3Do weakly-supervised disentangled representations generalize and prove useful for downstream tasks under covariate shifts and fairness constraints?
- RQ4How do adaptive weakly-supervised methods compare to fully supervised group-based methods in terms of performance and robustness?
Key findings
- Weakly-supervised models consistently outperform unsupervised baselines across five datasets.
- Ada-GVAE (and Ada-ML-VAE) reliably adapt to varying numbers of changed factors k and often match or exceed group-supervised methods.
- Disentangled representations learned under weak supervision correlate with strong generalization under covariate shifts.
- Weakly-supervised reconstruction loss serves as a useful proxy for downstream task performance and fairness outcomes.
- Representations learned via weak supervision improve robustness and abstract reasoning performance on downstream tasks.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.