Skip to main content
QUICK REVIEW

[Paper Review] Skewed link regression models for imbalanced binary response with applications to life insurance

Yin Shuang, Dipak K. Dey|arXiv (Cornell University)|Jul 30, 2020
Insurance, Mortality, Demography, Risk Management28 references4 citations
TL;DR

This paper proposes skewed link regression models—Generalized Extreme Value (GEV), Weibull, and Fréchet—using a fully Bayesian framework to address imbalanced binary response in life insurance mortality data, where death events are rare (e.g., ~1.3% annual rate). The models outperform standard logistic and probit regressions in predictive accuracy and model fit, especially under high skewness, as validated via real data and DIC-based model comparison.

ABSTRACT

For a portfolio of life insurance policies observed for a stated period of time, e.g., one year, mortality is typically a rare event. When we examine the outcome of dying or not from such portfolios, we have an imbalanced binary response. The popular logistic and probit regression models can be inappropriate for imbalanced binary response as model estimates may be biased, and if not addressed properly, it can lead to serious adverse predictions. In this paper, we propose the use of skewed link regression models (Generalized Extreme Value, Weibull, and Frechet link models) as more superior models to handle imbalanced binary response. We adopt a fully Bayesian approach for the generalized linear models (GLMs) under the proposed link functions to help better explain the high skewness. To calibrate our proposed Bayesian models, we use a real dataset of death claims experience drawn from a life insurance company's portfolio. Bayesian estimates of parameters were obtained using the Metropolis-Hastings algorithm and for Bayesian model selection and comparison, the Deviance Information Criterion (DIC) statistic has been used. For our mortality dataset, we find that these skewed link models are more superior than the widely used binary models with standard link functions. We evaluate the predictive power of the different underlying models by measuring and comparing aggregated death counts and death benefits.

Motivation & Objective

  • To address the limitations of symmetric link functions (logit, probit) in modeling highly imbalanced binary outcomes common in life insurance mortality data.
  • To develop flexible skewed link models capable of capturing extreme skewness in rare death events, such as annual mortality rates of ~1.3%.
  • To apply a fully Bayesian approach with MCMC (Metropolis-Hastings) for parameter estimation and model selection using DIC.
  • To evaluate predictive performance through aggregated death counts and death benefits using real insurance claims data.
  • To provide a robust framework for frequent mortality monitoring (e.g., quarterly) beyond annual tracking.

Proposed method

  • Proposes three skewed link functions—GEV, Weibull, and Fréchet—derived from extreme value distributions to model the cumulative distribution function of the latent variable in binary regression.
  • Adopts a fully Bayesian GLM framework with non-informative priors of the form π(α) ∝ 1/α^c (c > 1) to ensure proper posterior distributions.
  • Uses the Metropolis-Hastings algorithm for MCMC sampling to estimate posterior distributions of regression coefficients β and shape parameters α.
  • Applies the Deviance Information Criterion (DIC) for Bayesian model comparison and selection among competing link functions.
  • Employs latent variable representation z_i = x_i^T β + u_i, where u_i follows a Fréchet, Weibull, or GEV distribution, enabling flexible modeling of skewness.
  • Validates model propriety via theoretical proof showing that the posterior distribution is proper under full-rank design matrix and appropriate prior hyperparameters.

Experimental results

Research questions

  • RQ1Can skewed link models (GEV, Weibull, Fréchet) better capture the high skewness in imbalanced binary mortality outcomes than symmetric link functions?
  • RQ2Does a fully Bayesian approach with MCMC and DIC-based model selection improve parameter estimation and prediction accuracy in rare-event mortality modeling?
  • RQ3How do the proposed models compare in predictive performance when measuring aggregated death counts and total death benefits?
  • RQ4Can these models be effectively applied to more frequent mortality monitoring (e.g., quarterly) rather than annual assessments?
  • RQ5What are the theoretical conditions ensuring posterior propriety in the Fréchet link model under non-informative priors?

Key findings

  • The three proposed skewed link models—GEV, Weibull, and Fréchet—demonstrated superior performance over standard logistic, probit, and cloglog models in fitting the imbalanced mortality data.
  • The Bayesian approach with Metropolis-Hastings sampling successfully produced stable posterior estimates, and the DIC values indicated better model fit for the skewed link models.
  • The Fréchet link model was theoretically proven to have a proper posterior distribution under the specified non-informative prior π(α) ∝ 1/α^c with c > 1.
  • All three skewed link models showed improved predictive accuracy in estimating both aggregated death counts and total death benefits compared to conventional models.
  • The models are well-suited for frequent monitoring (e.g., quarterly) of mortality experience, enabling insurers to detect meaningful deviations from expected claims earlier.
  • The study confirms that symmetric link functions are inadequate for rare events with high skewness, and flexible skewed alternatives are essential for accurate actuarial modeling.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.