[Paper Review] Optimal Bayesian estimation of Gaussian mixtures with growing number of components
This paper develops a Bayesian framework for estimating Gaussian finite mixtures with a growing number of components, using a sample-size-dependent prior to achieve optimal posterior contraction rates under the Wasserstein distance. It establishes theoretical consistency in component estimation and provides a practical recipe for using Dirichlet process mixtures as a proxy for finite mixtures with adaptive component selection.
We study Bayesian estimation of finite mixture models in a general setup where the number of components is unknown and allowed to grow with the sample size. An assumption on growing number of components is a natural one as the degree of heterogeneity present in the sample can grow and new components can arise as sample size increases, allowing full flexibility in modeling the complexity of data. This however will lead to a high-dimensional model which poses great challenges for estimation. We novelly employ the idea of a sample size dependent prior in a Bayesian model and establish a number of important theoretical results. We first show that under mild conditions on the prior, the posterior distribution concentrates around the true mixing distribution at a near optimal rate with respect to the Wasserstein distance. Under a separation condition on the true mixing distribution, we further show that a better and adaptive convergence rate can be achieved, and the number of components can be consistently estimated. Furthermore, we derive optimal convergence rates for the higher-order mixture models where the number of components diverges arbitrarily fast. In addition, we suggest a simple recipe for using Dirichlet process (DP) mixture prior for estimating the finite mixture models and provide theoretical guarantees. In particular, we provide a novel solution for adopting the number of clusters in a DP mixture model as an estimate of the number of components in a finite mixture model. Simulation study and real data applications are carried out demonstrating the utilities of our method.
Motivation & Objective
- To address the gap in Bayesian estimation of finite mixture models where the number of components is unknown and allowed to grow with sample size.
- To establish posterior contraction rates for the mixing distribution under the Wasserstein distance in a high-dimensional, growing-component setting.
- To achieve adaptive and consistent estimation of the true number of components under a separation condition on the true mixing distribution.
- To provide a theoretically grounded method for using Dirichlet process mixture priors to estimate the number of components in finite mixtures.
- To demonstrate the practical utility of the method through simulations and real data applications on galaxy and geyser eruption data.
Proposed method
- Introduces a sample size-dependent prior distribution for the number of components in a finite mixture model, allowing the component count to grow with the sample size.
- Employs a Dirichlet process mixture (DPM) prior with a concentration parameter scaled by sample size to ensure consistency and adaptivity.
- Uses the posterior distribution of the number of clusters in a DPM model as a proxy estimate for the number of components in a finite mixture model.
- Applies the Wasserstein distance to measure convergence of the posterior to the true mixing distribution, enabling non-asymptotic theoretical analysis.
- Derives posterior contraction rates under mild regularity conditions, showing near-optimality in the minimax sense.
- Establishes theoretical guarantees for component selection consistency under a separation condition on the true mixing distribution.
Experimental results
Research questions
- RQ1Can Bayesian estimation of finite Gaussian mixtures achieve optimal posterior contraction rates when the number of components grows with the sample size?
- RQ2Under what conditions can the number of components be consistently estimated in a Bayesian framework with growing component count?
- RQ3How can a Dirichlet process mixture prior be used to estimate the number of components in a finite mixture model?
- RQ4What is the impact of hyperparameter choice (e.g., concentration parameter) on posterior inference for component count and density estimation?
- RQ5Can the posterior distribution of cluster counts in a DPM model reliably estimate the true number of components in a finite mixture?
Key findings
- Under mild conditions on the prior, the posterior distribution contracts around the true mixing distribution at a near-optimal rate with respect to the Wasserstein distance.
- Under a separation condition on the true mixing distribution, the posterior achieves a faster, adaptive convergence rate, and the number of components is consistently estimated.
- The minimax optimal convergence rate for mixing distribution estimation is of order $ n^{-1/(4(k^{ullet}-k_0)+2)} $, where $ k^{ullet} $ is the true number of components and $ k_0 $ is the number of well-separated components.
- The posterior distribution of the number of clusters in a Dirichlet process mixture model with a small concentration parameter concentrates near the true number of components, validating its use as a proxy for finite mixture estimation.
- Simulation results on galaxy and geyser data show that the choice of hyperparameters (e.g., Poisson mean or concentration parameter) significantly affects posterior inference, but the method remains robust and adaptive.
- The proposed method achieves optimal convergence rates even when the number of components diverges arbitrarily fast, demonstrating strong theoretical and practical flexibility.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.