[Paper Review] Compositional Data Regression in Insurance with Exponential Family PCA
This paper proposes exponential family principal component analysis (EPCA) to model compositional data in insurance, addressing limitations of traditional PCA when analyzing constrained, relative-proportion data. EPCA outperforms standard PCA and ILR-based methods by generating significant, interpretable principal components that enhance negative binomial regression accuracy for injury prediction in mining data.
Compositional data are multivariate observations that carry only relative information between components. Applying standard multivariate statistical methodology directly to analyze compositional data can lead to paradoxes and misinterpretations. Compositional data also frequently appear in insurance, especially with telematics information. However, such type of data does not receive deserved special treatment in most existing actuarial literature. In this paper, we explore and investigate the use of exponential family principal component analysis (EPCA) to analyze compositional data in insurance. The method is applied to analyze a dataset obtained from the U.S. Mine Safety and Health Administration. The numerical results show that EPCA is able to produce principal components that are significant predictors and improve the prediction accuracy of the regression model. The EPCA method can be a promising useful tool for actuaries to analyze compositional data.
Motivation & Objective
- Address the lack of specialized treatment for compositional data in actuarial science, particularly in telematics-driven insurance analytics.
- Overcome the pitfalls of standard multivariate methods on compositional data, such as scale invariance violations and Simpson’s paradox.
- Develop a robust dimensionality reduction technique tailored for compositional predictors with zero-inflated or sparse structures common in insurance data.
- Improve prediction accuracy in regression models by leveraging low-dimensional representations derived from compositional data.
- Demonstrate the superiority of EPCA over traditional PCA and ILR-based approaches in a real-world mining injury dataset.
Proposed method
- Apply exponential family principal component analysis (EPCA) to transform compositional predictors into orthogonal, low-dimensional components that preserve relative information.
- Use the log-likelihood of the exponential family distribution to model the data, enabling flexible handling of non-Gaussian, count-based, or skewed compositional data.
- Integrate EPCA with negative binomial regression to model overdispersed count outcomes (e.g., injury frequencies) while accounting for compositional covariates.
- Compare EPCA with classical PCA and isometric log-ratio (ILR) transformation methods in terms of model fit, significance, and prediction accuracy.
- Utilize principal components as predictors in regression models, where each component captures the dominant variation in the compositional structure.
- Apply inverse transformations to interpret EPCA results in the original compositional space, ensuring practical relevance for actuaries.
Experimental results
Research questions
- RQ1Can EPCA effectively extract meaningful, low-dimensional representations from compositional insurance data with zero-inflated components?
- RQ2How does EPCA compare to classical PCA and ILR-based methods in terms of statistical significance and predictive performance for regression models?
- RQ3Does the use of EPCA components improve the accuracy of negative binomial regression models when predicting injury frequencies in mining operations?
- RQ4To what extent do EPCA components preserve subcomposition coherence and avoid paradoxes like Simpson’s paradox in compositional data analysis?
- RQ5Can EPCA handle sparse compositional data common in telematics-based insurance data without requiring data transformation that excludes zeros?
Key findings
- EPCA-generated principal components were all statistically significant (p < 0.05), whereas the third principal component from classical PCA was not (p = 0.8214).
- The NBEPCA model (using EPCA components) achieved higher predictive accuracy than the NBPCA and ILR-based models, as evidenced by improved model fit and significance of components.
- In the ILR-transformed model, several components (e.g., V5, V7) had p-values above 0.1, indicating they were not significant, suggesting the need for dimensionality reduction.
- The EPCA method successfully handled zero-inflated compositional data without requiring data transformation that excludes zero values.
- The EPCA components were more interpretable and stable than ILR-transformed variables, which showed highly variable and difficult-to-interpret coefficients.
- The NBEPCA model demonstrated superior performance with a deviance reduction and better significance of predictors, confirming EPCA’s value in insurance regression.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.