[Paper Review] Parsimonious mixtures of contaminated Gaussian distributions with application to allometric studies
This paper proposes a parsimonious finite mixture model based on contaminated Gaussian distributions for robust model-based clustering, incorporating automatic estimation of outlier proportions and contamination levels without prior specification. The method uses eigen-decomposition for covariance structure parsimony and demonstrates strong performance in allometric studies and simulations compared to standard elliptical mixtures.
A mixture of contaminated Gaussian distributions is developed for model-based clustering. In addition to the parameters of the classical Gaussian mixture, each component of our contaminated mixture has a parameter controlling the proportion of outliers, spurious points, or noise (collectively referred to as bad points herein) and one specifying the degree of contamination. Crucially, these parameters do not have to be specified a priori, adding a flexibility to our approach. Parsimony is introduced via eigen-decomposition of the component covariance matrices, and sufficient conditions for the identifiability of all the members of the resulting family are provided. An expectation-conditional maximization algorithm is outlined for parameter estimation and various implementation issues are discussed. Using a large scale simulation study, we investigate the behavior of the proposed approach and we provide a comparison with finite mixture models of some well-established multivariate elliptical distributions. The performance of this novel family of models is also illustrated on artificial and real data, with particular emphasis to the application in allometric studies.
Motivation & Objective
- To develop a flexible finite mixture model that explicitly accounts for outliers and noisy data in clustering applications.
- To allow automatic estimation of contamination parameters (outlier proportion and degree) without requiring prior specification.
- To ensure model identifiability through sufficient conditions on the component parameters.
- To apply the model to allometric studies, where data often contain measurement errors and extreme values.
- To compare performance against established multivariate elliptical mixture models using simulation and real data.
Proposed method
- The model extends classical Gaussian finite mixtures by adding two component-specific parameters: one for the proportion of bad points (outliers/noise), and one for the degree of contamination.
- Parsimony is achieved via eigen-decomposition of component covariance matrices, reducing the number of free parameters.
- An expectation-conditional maximization (ECM) algorithm is used for parameter estimation, enabling iterative optimization of the likelihood.
- Sufficient conditions for identifiability of all components in the mixture family are formally derived and provided.
- The approach is implemented with attention to numerical stability and convergence in high-dimensional settings.
- Model fitting is evaluated using likelihood-based criteria and compared to standard elliptical mixture models.
Experimental results
Research questions
- RQ1How can a finite mixture model be made robust to outliers and noise without requiring prior knowledge of contamination levels?
- RQ2What conditions ensure the identifiability of components in a contaminated Gaussian mixture with parsimonious covariance structures?
- RQ3How does the proposed model perform in clustering tasks when data contain outliers or measurement errors, especially in allometric studies?
- RQ4How does the performance of the contaminated Gaussian mixture compare to standard elliptical mixture models in simulation and real data scenarios?
- RQ5Can the automatic estimation of contamination parameters improve clustering accuracy and robustness in practice?
Key findings
- The proposed contaminated Gaussian mixture model successfully estimates both the proportion of bad points and the degree of contamination without requiring prior specification.
- The use of eigen-decomposition enables parsimonious modeling of covariance matrices while maintaining model flexibility and interpretability.
- Sufficient conditions for identifiability of all components in the mixture family are established, ensuring statistical validity.
- Simulation results show the model outperforms standard finite mixtures of elliptical distributions in the presence of contamination and outliers.
- The model demonstrates strong empirical performance on real allometric data, effectively handling measurement errors and extreme observations.
- The ECM algorithm converges reliably and provides stable parameter estimates across diverse data configurations.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.