[Paper Review] Semiparametric Mixture Regression with Unspecified Error Distributions
This paper proposes a semiparametric mixture of linear regression models that relax the normality assumption on error distributions by allowing unspecified component error densities. Using a kernel density estimation-based EM algorithm, the method achieves consistent estimation and asymptotic normality, significantly outperforming traditional MLE under non-normal errors while maintaining comparable performance under normality, as demonstrated in simulations and a real equine viral load dataset.
In fitting a mixture of linear regression models, normal assumption is traditionally used to model the error and then regression parameters are estimated by the maximum likelihood estimators (MLE). This procedure is not valid if the normal assumption is violated. To relax the normal assumption on the error distribution hence reduce the modeling bias, we propose semiparametric mixture of linear regression models with unspecified error distributions. We establish a more general identifiability result under weaker conditions than existing results, construct a class of new estimators, and establish their asymptotic properties. These asymptotic results also apply to many existing semiparametric mixture regression estimators whose asymptotic properties have remained unknown due to the inherent difficulties in obtaining them. Using simulation studies, we demonstrate the superiority of the proposed estimators over the MLE when the normal error assumption is violated and the comparability when the error is normal. Analysis of a newly collected Equine Infectious Anemia Virus data in 2017 is employed to illustrate the usefulness of the new estimator.
Motivation & Objective
- To address the modeling bias introduced by assuming normal errors in traditional mixture regression models.
- To develop a semiparametric estimator that allows unspecified error densities while maintaining identifiability and asymptotic properties.
- To establish the consistency and asymptotic normality of the proposed estimator under weaker conditions than existing methods.
- To demonstrate improved classification accuracy and robustness in finite samples when errors deviate from normality.
- To provide a theoretically grounded, flexible alternative to parametric MLE in mixture regression settings.
Proposed method
- Proposes a semiparametric mixture regression model where component error densities are unspecified, replacing the normal density with a general density function g with mean zero and unit variance.
- Introduces a semiparametric EM algorithm that alternates between estimating regression parameters θ and nonparametrically estimating the error density g using kernel density estimation.
- Uses a kernel density estimator for the error density with a bandwidth selected via cross-validation to ensure smoothness and consistency.
- Establishes identifiability of the model under general conditions, including non-identical scale parameters τj and arbitrary component densities g_j.
- Derives the asymptotic distribution of the proposed estimator, showing consistency and asymptotic normality under regularity conditions.
- Applies the method to real data using leave-one-out cross-validation to assess classification performance.
Experimental results
Research questions
- RQ1Can a semiparametric mixture regression model with unspecified error densities achieve consistent estimation without assuming normality?
- RQ2What are the minimal regularity conditions required for model identifiability in a general semiparametric mixture regression framework?
- RQ3How does the performance of the proposed estimator compare to MLE when the error distribution is non-normal?
- RQ4Does the proposed method improve classification accuracy in finite samples compared to traditional MLE?
- RQ5Can the asymptotic theory of the proposed estimator be applied to existing semiparametric mixture regression estimators whose asymptotic properties were previously unknown?
Key findings
- The proposed KDEEM estimator achieved 100% correct classification percentage (CCP) on the EIAV dataset, compared to 93.33% for MLEEM, demonstrating superior classification performance.
- Under non-normal error distributions, the proposed estimator significantly outperformed MLE in simulation studies, reducing modeling bias and improving estimation accuracy.
- The method maintained comparable performance to MLE when errors were actually normal, indicating robustness across distributional assumptions.
- Theoretical results establish the consistency and asymptotic normality of the proposed estimator under weaker conditions than previous identifiability results.
- The asymptotic theory applies to many existing semiparametric mixture regression estimators whose asymptotic properties were previously unproven.
- The kernel-based estimation of the error density allowed flexible modeling of complex error structures without parametric assumptions.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.