[Paper Review] Infinite Mixtures of Infinite Factor Analysers: Nonparametric Model-Based Clustering via Latent Gaussian Models
This paper proposes Infinite Mixtures of Infinite Factor Analysers (IMIFA), a nonparametric Bayesian model that jointly infers the number of clusters and cluster-specific numbers of latent factors in high-dimensional data using shrinkage priors and Dirichlet processes. The method enables automatic, adaptive clustering without pre-specifying model dimensions, using a stick-breaking construction and slice sampling for efficient posterior inference.
Gaussian mixture models for high-dimensional data often assume a factor analytic covariance structure within each mixture component. When clustering via such mixtures of factor analysers (MFA), the numbers of clusters and latent factors must be specified in advance of model fitting, and remain fixed. The pair which optimises some model selection criterion are usually chosen. Within such a computationally intensive model search, only models in which the number of factors is common across clusters are generally considered. Here the mixture of infinite factor analysers (MIFA) is introduced, which allows different clusters to have different numbers of factors through the use of shrinkage priors. The cluster-specific number of factors is automatically inferred during model fitting, via an efficient, adaptive Gibbs sampler. However, the number of clusters still requires selection. Infinite mixtures of infinite factor analysers (IMIFA) is a nonparameteric extension of the MIFA model which uses Dirichlet or Poisson-Dirichlet processes to facilitate automatic, simultaneous inference of the number of clusters and the cluster-specific number of factors. IMIFA provides a flexible approach to fitting mixtures of factor analysers which obviates the need for model selection criteria. Estimation uses the stick-breaking construction and a slice sampler. Application to simulated data, a benchmark dataset, and spectral metabolomic data illustrate the methodology and its performance.
Motivation & Objective
- To address the limitation of traditional mixtures of factor analysers (MFA) that require pre-specifying both the number of clusters and the number of factors per cluster.
- To develop a flexible, nonparametric model that allows cluster-specific numbers of latent factors through shrinkage priors.
- To eliminate the need for model selection criteria by enabling simultaneous automatic inference of both the number of clusters and the cluster-specific factor dimensions.
- To improve computational efficiency and scalability in high-dimensional clustering tasks using adaptive Gibbs sampling and stick-breaking constructions.
Proposed method
- Introduces the MIFA model as a foundation, using shrinkage priors to allow different clusters to have varying numbers of latent factors.
- Extends MIFA to IMIFA by employing a Dirichlet or Poisson-Dirichlet process to nonparametrically infer the number of clusters.
- Employs a stick-breaking construction to represent the infinite mixture of components, enabling flexible clustering without prespecifying the number of clusters.
- Uses a slice sampler for posterior computation, allowing efficient exploration of the infinite-dimensional space of cluster and factor configurations.
- Applies shrinkage priors to the factor loading matrices to encourage sparsity and automatically determine the effective number of factors per cluster.
- Develops an adaptive Gibbs sampler that jointly updates cluster assignments, factor dimensions, and model parameters in a single MCMC framework.
Experimental results
Research questions
- RQ1Can a nonparametric Bayesian model jointly infer the number of clusters and the cluster-specific number of latent factors in high-dimensional data?
- RQ2How does the use of shrinkage priors and Dirichlet processes improve model flexibility compared to fixed-factor MFA models?
- RQ3To what extent can IMIFA reduce the need for model selection criteria in clustering applications?
- RQ4How does the performance of IMIFA compare to existing methods on simulated data, benchmark datasets, and real metabolomic data?
Key findings
- IMIFA successfully infers both the number of clusters and cluster-specific numbers of factors without requiring pre-specification or model selection criteria.
- The use of shrinkage priors enables automatic determination of the effective number of factors per cluster, with some clusters having fewer factors than others.
- The stick-breaking construction and slice sampler allow efficient posterior exploration of the infinite mixture space, supporting scalable inference.
- On simulated data, IMIFA accurately recovered the true number of clusters and factor dimensions, outperforming fixed-factor MFA models.
- On a benchmark dataset, IMIFA identified a more parsimonious and interpretable clustering structure than competing methods.
- In spectral metabolomic data, IMIFA revealed biologically meaningful clusters with varying latent dimensionality, suggesting heterogeneous biological substructures.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.