[Paper Review] How to Not Measure Disentanglement
This paper identifies critical flaws in existing disentanglement metrics, showing they fail to consistently assign high scores to truly disentangled representations and low scores to entangled ones. It proposes a new metric, 3CharM, which theoretically satisfies both desirable properties and unifies the strengths of prior metrics through a formalized framework grounded in disentanglement characteristics.
To evaluate disentangled representations several metrics have been proposed. However, theoretical guarantees for conventional metrics of disentanglement are missing. Moreover, conventional metrics do not have a consistent correlation with the outcomes of qualitative studies. In this paper we analyze metrics of disentanglement and their properties. We conclude that existing metrics of disentanglement were created to reflect different characteristics of disentanglement and do not satisfy two basic desirable properties: (1) assign a high score to representations that are disentangled according to the definition; and (2) assign a low score to representations that are entangled according to the definition. In addition, we propose a new metric of disentanglement and prove that it satisfies both of the properties.
Motivation & Objective
- To analyze why conventional disentanglement metrics fail to correlate with qualitative assessments of disentanglement.
- To identify and formalize two fundamental characteristics that define disentangled representations: (1) one latent factor per generative factor, and (2) invertible mapping between latent and generative factors.
- To demonstrate that existing metrics like BetaVAE, FactorVAE, DCI, SAP, and MIG do not consistently satisfy basic desirable properties: high scores for disentangled representations and low scores for entangled ones.
- To propose a new metric, 3CharM, that theoretically satisfies both desirable properties and unifies the evaluation of disentanglement across different characteristics.
Proposed method
- The authors define disentanglement based on two core characteristics: (1) each generative factor should be encoded by a single, independent latent factor, and (2) the mapping between generative and latent factors should be invertible.
- They analyze existing metrics (BetaVAE, FactorVAE, DCI, SAP, MIG) by evaluating whether they assign high scores to representations satisfying the defined characteristics and low scores to those that do not.
- The paper introduces 3CharM, a new metric that combines the evaluation of both disentanglement characteristics into a single, theoretically grounded measure using mutual information and invertibility constraints.
- 3CharM is constructed by measuring the degree to which each latent factor captures exactly one generative factor, while ensuring that the representation remains invertible under the given constraints.
- Theoretical proofs are provided to show that 3CharM satisfies both desirable properties: it scores high on disentangled representations and low on entangled ones.
Experimental results
Research questions
- RQ1Why do existing disentanglement metrics fail to correlate with qualitative assessments of disentanglement?
- RQ2What fundamental properties should a reliable disentanglement metric satisfy?
- RQ3Do current metrics like BetaVAE, FactorVAE, DCI, SAP, and MIG satisfy the basic desirable properties of assigning high scores to disentangled representations and low scores to entangled ones?
- RQ4Can a unified metric be designed that reflects both key characteristics of disentanglement and is theoretically justified?
Key findings
- Most conventional disentanglement metrics—such as BetaVAE, FactorVAE, DCI, SAP, and MIG—fail to assign high scores to representations that are truly disentangled according to the definition.
- These metrics also fail to assign low scores to entangled representations, indicating a lack of consistency with the theoretical definition of disentanglement.
- The study reveals that metrics are designed to reflect different characteristics of disentanglement, leading to inconsistent rankings across datasets and models.
- The proposed 3CharM metric is proven to satisfy both desirable properties: it scores highly on disentangled representations and low on entangled ones.
- 3CharM unifies the evaluation of disentanglement by incorporating both the one-to-one correspondence and invertibility between generative and latent factors.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.