Skip to main content
QUICK REVIEW

[論文レビュー] A Fair Pricing Model via Adversarial Learning

Vincent Grari, Arthur Charpentier|arXiv (Cornell University)|Feb 24, 2022
Insurance, Mortality, Demography, Risk Management被引用数 11
ひとこと要約

本稿では、地理的要因と車両関連要因を分離することでバイアスを低減しつつ予測精度を維持する、フェアな保険料設定のための新しい敵対的オートエンコーダーに基づくフレームワークを提案する。統合的なモデルを用いてバイアスのない保険料構成要素を生成することで、合成データおよび実世界のデータセットの両方で、従来の手法よりも公平性と性能の両面で優れている。

ABSTRACT

At the core of insurance business lies classification between risky and non-risky insureds, actuarial fairness meaning that risky insureds should contribute more and pay a higher premium than non-risky or less-risky ones. Actuaries, therefore, use econometric or machine learning techniques to classify, but the distinction between a fair actuarial classification and "discrimination" is subtle. For this reason, there is a growing interest about fairness and discrimination in the actuarial community Lindholm, Richman, Tsanakas, and Wuthrich (2022). Presumably, non-sensitive characteristics can serve as substitutes or proxies for protected attributes. For example, the color and model of a car, combined with the driver's occupation, may lead to an undesirable gender bias in the prediction of car insurance prices. Surprisingly, we will show that debiasing the predictor alone may be insufficient to maintain adequate accuracy (1). Indeed, the traditional pricing model is currently built in a two-stage structure that considers many potentially biased components such as car or geographic risks. We will show that this traditional structure has significant limitations in achieving fairness. For this reason, we have developed a novel pricing model approach. Recently some approaches have Blier-Wong, Cossette, Lamontagne, and Marceau (2021); Wuthrich and Merz (2021) shown the value of autoencoders in pricing. In this paper, we will show that (2) this can be generalized to multiple pricing factors (geographic, car type), (3) it perfectly adapted for a fairness context (since it allows to debias the set of pricing components): We extend this main idea to a general framework in which a single whole pricing model is trained by generating the geographic and car pricing components needed to predict the pure premium while mitigating the unwanted bias according to the desired metric.

研究の動機と目的

  • 規制的・倫理的懸念が高まる中、保険料設定における公平性の増大するニーズに対応すること。
  • 非感受性属性(例:場所、車両タイプ)からのプロキシバイアスを緩和するが、予測精度を損なわない手法を開発すること。
  • オートエンコーダーに基づく表現学習を、公平性制約を設けた統合的かつエンドツーエンドの保険料設定モデルへと拡張すること。
  • 敵対的デバイアスが、保険料設定の文脈において、標準的なフェアML手法よりも効果的であることを示すこと。
  • 連続的およびカテゴリカルなリスク要因を有する実世界の保険データに適用可能な汎用的フレームワークを提供すること。

提案手法

  • フレームワークは、入力特徴から地理的および車両関連リスク要因の分離表現を学習するためのディープオートエンコーダーを用いる。
  • 学習されたリスク成分と保護された属性(例:性別、場所によるレースのプロキシ)の依存関係を最小化する敵対的トレーニング目的関数を導入する。
  • ハイパーパrameter λ が、公平性(最大相関係数で測定)と純保険料の予測精度のトレードオフを制御する。
  • 純保険料予測と敵対的デバイアスの両方を同時に最適化することで、二段階処理を必要としないエンドツーエンドのトレーニングが可能になる。
  • 本手法は複数の定価要因に一般化可能であり、離散的および連続的両方の感受性属性に対応可能で、空間データをレースのプロキシとして使用可能である。
  • 再構成損失、予測損失、敵対的損失を組み合わせた修正された損失関数を用い、公平性と精度のバランスをとる。

実験結果

リサーチクエスチョン

  • RQ1敵対的オートエンコーダーは、保険料設定のためのフェアで分離されたリスク要因を効果的に生成するために適応可能か?
  • RQ2提案されたフレームワークは、地理的および車両関連要因からのプロキシバイアスを低減しつつ、予測精度を維持または向上させられるか?
  • RQ3ハイパーパrameter λ の変化に伴い、公平性(最大相関係数を用いて測定)とモデル精度のトレードオフはどのように変化するか?
  • RQ4本手法は、保険料設定の文脈において、標準的なフェアMLアルゴリズムよりも効果的か?
  • RQ5本モデルは、複雑で高次元のリスク要因と空間的プロキシを有する実世界のデータセットにも一般化可能か?

主な発見

  • 提案された AE_HGR モデルは、感受性属性と予測との間の依存関係を顕著に低減しており、λ の値が高いほどフランスの地方ごとに滑らかでより公平な予測が得られる。
  • 空間的バイアス低減タスクにおいて、λ を 0 から 0.6 に増加させることで、44番目の地方に属する高リスクの INSEE コード 29830 において、極端な事故発生確率の予測(例:0.8 以上から約 0.4~0.6 に)が低下した。
  • ベースラインのGLMおよび標準的なフェアML手法と比較して、本モデルは公平性と精度のトレードオフをより優れたバランスで達成しており、特にバイアス低減の際の性能維持が顕著に優れている。
  • オートエンコーダーに基づく地理的および車両リスク要因の分離により、従来の二段階定価モデルと比較して、全体の予測精度が向上した。
  • 感受性属性が直接利用可能でない場合でも、場所を保護された特性のプロキシとして活用することで、空間的バイアスを効果的に中和した。
  • 実証的結果から、本モデルの公平性向上が、ナーブなデバイアス手法とは異なり、重要な予測信号を破棄することなく達成されたことが確認された。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。