[论文解读] Mitigating Discrimination in Insurance with Wasserstein Barycenters
本文提出使用Wasserstein中位数来减轻保险定价中的歧视问题,通过一种保留分布差异而非依赖简单均值缩放的方式,在不同人口群体间调整风险预测。该方法在真实机动车保险数据上验证,显著减少了人口平等性违规,同时保持了预测准确性,有效降低了保费中的不公平差异。
The insurance industry is heavily reliant on predictions of risks based on characteristics of potential customers. Although the use of said models is common, researchers have long pointed out that such practices perpetuate discrimination based on sensitive features such as gender or race. Given that such discrimination can often be attributed to historical data biases, an elimination or at least mitigation is desirable. With the shift from more traditional models to machine-learning based predictions, calls for greater mitigation have grown anew, as simply excluding sensitive variables in the pricing process can be shown to be ineffective. In this article, we first investigate why predictions are a necessity within the industry and why correcting biases is not as straightforward as simply identifying a sensitive variable. We then propose to ease the biases through the use of Wasserstein barycenters instead of simple scaling. To demonstrate the effects and effectiveness of the approach we employ it on real data and discuss its implications.
研究动机与目标
- 解决由与年龄、种族或性别等敏感属性相关的历史数据偏差所导致的不公平保险定价这一长期问题。
- 克服在基于机器学习的风险模型中,通过简单变量排除或基于均值的缩放方法实现公平性的局限性。
- 开发一种方法,通过考虑风险评分的完整分布而非仅其均值,公平地调整各群体的预测。
- 提供一种理论基础坚实、可解释且在实证中有效的公平性方法,应用于精算科学中的最优传输理论。
- 证明Wasserstein中位数可在保持模型性能的同时减少现实世界保险数据集中的群体平等性违规。
提出的方法
- 该方法利用最优传输理论,计算不同敏感群体风险预测分布的Wasserstein中位数,实现预测的公平调整。
- 它构建一个参考分布(中位数),使该分布到各群体特定预测分布的Wasserstein距离最小化,从而确保最小失真。
- 该方法将简单的缩放(例如,将预测乘以一个因子)替换为考虑各群体风险分布形状与离散度的分布感知变换。
- 中位数通过凸优化框架计算,从而实现公平预测的高效且稳定估计。
- 该方法应用于以年龄为敏感属性的真实机动车保险数据,对比了GLM、GBM和RF模型的性能。
- 公平性通过人口平等性进行评估,结果表明该方法在不牺牲预测准确性的情况下减少了不公平差异。

实验结果
研究问题
- RQ1与传统缩放方法相比,Wasserstein中位数是否能有效减少保险定价中的群体平等性违规?
- RQ2通过最优传输实现的分布调整与基于均值的缩放相比,在保持公平性与预测性能方面表现如何?
- RQ3当敏感属性未在建模中直接使用时,Wasserstein中位数在多大程度上可缓解间接歧视?
- RQ4该方法在不同机器学习模型(GLM、GBM、RF)下,对真实保险数据的可解释性与鲁棒性如何?
- RQ5基于中位数的调整对年轻与年长驾驶员之间风险预测差异有何影响?
主要发现
- 对于年长驾驶员(年龄 >65 岁),中位数方法将预测索赔概率的差异从基线的5.27%(初始风险低于5%)降低至4.71%,显著提升了公平性。
- 当初始风险为20%时,基线模型预测年长驾驶员的索赔概率为21.26%,而中位数方法将其降低至18.85%,缩小了差距。
- 对于年轻驾驶员(年龄 <30 岁),当初始风险为20%时,该方法将预测风险从基线的14.84%降低至12.54%,表明有效缓解了过度惩罚。
- 中位数方法在减少群体平等性违规方面优于简单缩放,尤其在风险较高的群体中,差异最为显著。
- 该方法在各类模型(GLM、GBM、RF)中均保持了良好的预测准确性,显示出在真实保险环境中的鲁棒性与实际可行性。
- 匹配预测的可视化结果表明,与基于均值的缩放相比,中位数方法在风险分布的尾部区域更好地对齐了各群体的分布。
![Figure 2: Distributions of $m(\boldsymbol{x},s={\color[rgb]{0,0.62890625,0.54296875}\definecolor[named]{pgfstrokecolor}{rgb}{0,0.62890625,0.54296875}\text{{A}}})$ and $m(\boldsymbol{x},s={\color[rgb]{0.94921875,0.6796875,0}\definecolor[named]{pgfstrokecolor}{rgb}{0.94921875,0.6796875,0}\text{{B}}})$](https://ar5iv.labs.arxiv.org/html/2306.12912/assets/figs/gender/unnamed-chunk-3-1.png)
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。