[论文解读] Skewed link regression models for imbalanced binary response with applications to life insurance
本文提出了一类偏斜链接回归模型——广义极值(GEV)、威布尔和弗雷歇分布——采用完全贝叶斯框架,以解决人寿保险死亡数据中罕见死亡事件(如年死亡率约1.3%)导致的类别不平衡二值响应问题。在真实数据和基于DIC的模型比较下,这些模型在预测准确性和模型拟合度方面均优于标准的逻辑斯蒂和 probit 回归,尤其在高度偏斜的情况下表现更优。
For a portfolio of life insurance policies observed for a stated period of time, e.g., one year, mortality is typically a rare event. When we examine the outcome of dying or not from such portfolios, we have an imbalanced binary response. The popular logistic and probit regression models can be inappropriate for imbalanced binary response as model estimates may be biased, and if not addressed properly, it can lead to serious adverse predictions. In this paper, we propose the use of skewed link regression models (Generalized Extreme Value, Weibull, and Frechet link models) as more superior models to handle imbalanced binary response. We adopt a fully Bayesian approach for the generalized linear models (GLMs) under the proposed link functions to help better explain the high skewness. To calibrate our proposed Bayesian models, we use a real dataset of death claims experience drawn from a life insurance company's portfolio. Bayesian estimates of parameters were obtained using the Metropolis-Hastings algorithm and for Bayesian model selection and comparison, the Deviance Information Criterion (DIC) statistic has been used. For our mortality dataset, we find that these skewed link models are more superior than the widely used binary models with standard link functions. We evaluate the predictive power of the different underlying models by measuring and comparing aggregated death counts and death benefits.
研究动机与目标
- 解决对称链接函数(logit、probit)在建模人寿保险死亡数据中常见高度不平衡二值结果时的局限性。
- 开发能够捕捉罕见死亡事件极端偏斜性的灵活偏斜链接模型,例如年死亡率约为1.3%的情况。
- 采用完全贝叶斯方法,结合MCMC(Metropolis-Hastings)进行参数估计,并利用DIC进行模型选择。
- 通过真实保险索赔数据中的汇总死亡人数和死亡赔付总额,评估预测性能。
- 为频繁死亡率监测(例如每季度一次)提供稳健框架,超越传统的年度追踪。
提出的方法
- 提出三种源自极值分布的偏斜链接函数——GEV、Weibull 和 Fréchet——用于建模二值回归中潜变量的累积分布函数。
- 采用非信息性先验形式 π(α) ∝ 1/α^c(c > 1)的完全贝叶斯广义线性模型框架,以确保后验分布的合理性。
- 使用 Metropolis-Hastings 算法进行 MCMC 抽样,以估计回归系数 β 和形状参数 α 的后验分布。
- 应用偏差信息准则(DIC)进行贝叶斯模型比较与选择,评估不同链接函数之间的表现。
- 采用潜变量表示 z_i = x_i^T β + u_i,其中 u_i 服从 Fréchet、Weibull 或 GEV 分布,从而实现对偏斜性的灵活建模。
- 通过理论证明验证模型的合理性:在设计矩阵满秩且先验超参数适当时,后验分布为正则分布。
实验结果
研究问题
- RQ1与对称链接函数相比,偏斜链接模型(GEV、Weibull、Fréchet)是否能更好地捕捉类别不平衡二值死亡结果中的高度偏斜性?
- RQ2采用完全贝叶斯方法结合 MCMC 和基于 DIC 的模型选择,是否能提升罕见事件死亡率建模中的参数估计与预测准确性?
- RQ3在测量汇总死亡人数和总死亡赔付金额时,所提出的模型在预测性能上如何比较?
- RQ4这些模型是否可有效应用于更频繁的死亡率监测(如每季度一次),而非仅限于年度评估?
- RQ5在非信息性先验下,Fréchet 链接模型的后验分布正则性的理论条件是什么?
主要发现
- 所提出的三种偏斜链接模型——GEV、Weibull 和 Fréchet——在拟合类别不平衡死亡率数据方面,表现优于标准的逻辑斯蒂、probit 和 cloglog 模型。
- 结合 Metropolis-Hastings 抽样的贝叶斯方法成功生成了稳定的后验估计,DIC 值表明偏斜链接模型具有更优的模型拟合度。
- 理论证明表明,在指定的非信息性先验 π(α) ∝ 1/α^c(c > 1)下,Fréchet 链接模型的后验分布为正则分布。
- 所有三种偏斜链接模型在估计汇总死亡人数和总死亡赔付金额方面,均展现出优于传统模型的预测准确性。
- 这些模型非常适合用于频繁的死亡率监测(如每季度一次),使保险公司能够更早地检测到与预期索赔的显著偏差。
- 本研究证实,对称链接函数在高度偏斜的罕见事件中表现不足,而灵活的偏斜替代方案对精确精算建模至关重要。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。