[論文レビュー] On the combination of graph data for assessing thin-file borrowers' creditworthiness
本稿では、手作業で作成した特徴量、グラフ埋め込み(Node2Vec)、およびグラフニューラルネットワーク(GNNs)を組み合わせたハイブリッドフレームワークを提案し、信用履歴が乏しい借り手(個人および企業)の信用スコーリングを改善する。複数のグラフ表現学習(GRL)手法の出力を勾配ブースティング分類器に集約することで、個人および企業の信用スコーリングの両方で予測性能が著しく向上し、GNNsと特徴工学が強く相乗効果を発揮する一方、Node2Vecはほとんど寄与しない。
The thin-file borrowers are customers for whom a creditworthiness assessment is uncertain due to their lack of credit history; many researchers have used borrowers' relationships and interactions networks in the form of graphs as an alternative data source to address this. Incorporating network data is traditionally made by hand-crafted feature engineering, and lately, the graph neural network has emerged as an alternative, but it still does not improve over the traditional method's performance. Here we introduce a framework to improve credit scoring models by blending several Graph Representation Learning methods: feature engineering, graph embeddings, and graph neural networks. We stacked their outputs to produce a single score in this approach. We validated this framework using a unique multi-source dataset that characterizes the relationships and credit history for the entire population of a Latin American country, applying it to credit risk models, application, and behavior, targeting both individuals and companies. Our results show that the graph representation learning methods should be used as complements, and these should not be seen as self-sufficient methods as is currently done. In terms of AUC and KS, we enhance the statistical performance, outperforming traditional methods. In Corporate lending, where the gain is much higher, it confirms that evaluating an unbanked company cannot solely consider its features. The business ecosystem where these firms interact with their owners, suppliers, customers, and other companies provides novel knowledge that enables financial institutions to enhance their creditworthiness assessment. Our results let us know when and which group to use graph data and what effects on performance to expect. They also show the enormous value of graph data on the unbanked credit scoring problem, principally to help companies' banking.
研究の動機と目的
- 信用履歴が乏しい借り手(信用履歴が最小限または全くない個人および企業)の信用スコーリングの課題に対処すること。
- 単一のグラフ表現学習(GRL)手法に限界があるのを補うために、複数の手法を組み合わせること。
- 複数の情報源、全国規模のネットワークおよび金融データを用いて、個人および企業の信用スコーリングにおける予測性能を向上させること。
- ソーシャルネットワークデータが信用力評価において、いつ、どこで最も価値をもたらすかを特定すること。
- 説明可能なAI(SHAP)を用いて、特徴量の重要性およびモデルの挙動に関する実用的洞察を提供すること。
提案手法
- 手作業で作成した特徴量、Node2Vec埋め込み、およびグラフニューラルネットワーク(GNNs)の3つのGRL手法を統合する統一された情報処理フレームワークを提案する。
- 各GRL手法の予測結果を1つの入力ベクトルに集約し、勾配ブースティング分類器(XGBoost)に供給する。
- 金融取引、社会的関係、企業関係を含む、ラテンアメリカの1か国の全人口をカバーする独自の全国的データセットを用いる。
- 個人および企業の申請スコーリング、および行動スコーリングの4つの信用スコーリングシナリオに、このフレームワークを適用する。
- SHAP値を用いて特徴量の寄与度とモデルの意思決定を解釈し、どのネットワーク特徴量が予測に寄与しているかを明らかにする。
- 標準指標(受信者操作特性曲線下積分(AUC)およびコルモゴロフ=スミルノフ統計量(KS))を用いて性能を検証する。
実験結果
リサーチクエスチョン
- RQ1複数のグラフ表現学習(GRL)手法を組み合わせることで、単独で使用する場合と比較して、信用スコーリングの性能が向上するか?
- RQ2統合されたネットワーク特徴量は、信用リスクについてどのような洞察を提供し、意思決定をどのように向上させるか?
- RQ3個人または企業の信用スコーリングの文脈において、ソーシャルネットワークデータが最も高い性能向上をもたらすのはどちらか?
- RQ4どの種類のネットワーク特徴量(例:エゴネットワーク、家族関係、サプライヤー網)が信用力の予測に最も予測的か?
- RQ5異なるGRL手法(特徴工学、Node2Vec、GNNs)の寄与度は、借り手のタイプやスコーリングシナリオによってどのように変化するか?
主な発見
- 手作業で作成した特徴量とGNNsの組み合わせが、最も高い性能向上をもたらし、AUCおよびKS指標がベースラインモデルを著しく上回る。
- 未銀行取引者申請スコアリングにおいて、提案されたフレームワークが最大の性能向上を達成しており、金融的排除状態にある人々にとっての価値を示している。
- 企業融資において、経営者、サプライヤー、顧客を含むビジネスエコシステムが、企業固有の属性を超えた重要な予測信号を提供する。
- BenchScoreベースラインは依然として重要であるが、ネットワーク特徴量と組み合わせるとその相対的影響は低下し、ネットワークデータが従来の指標を補完するのではなく強化することを示している。
- Node2Vec埋め込みはモデル性能にほとんど寄与しないため、この文脈では単独での特徴工学手法としての有用性が限定的であると考えられる。
- 個人信用スコーリングにおいて、家族関係特徴量(FamilyNet)は広範な社会的ネットワーク(EOWNet)よりもより情報量が多く、個人の信用評価において血縁関係の重要性を示している。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。