[논문 리뷰] On the combination of graph data for assessing thin-file borrowers' creditworthiness
이 논문은 수작업으로 만든 특징, 그래프 임베딩(Node2Vec), 그래프 신경망(GNNs)을 조합한 하이브리드 프레임워크를 제안하여, 신용 기록이 부족한 개인 및 기업(이른바 '얇은 파일' 대상)의 신용 평가를 향상시킨다. 여러 그래프 표현 학습(GRL) 방법의 출력을 기울기 부스팅 분류기로 통합함으로써, 개인 및 기업의 신용 평가에서 예측 성능이 크게 향상되며, GNNs와 수작업 특징 간의 상호보완성이 두드러지지만, Node2Vec는 기여도가 미미하다.
The thin-file borrowers are customers for whom a creditworthiness assessment is uncertain due to their lack of credit history; many researchers have used borrowers' relationships and interactions networks in the form of graphs as an alternative data source to address this. Incorporating network data is traditionally made by hand-crafted feature engineering, and lately, the graph neural network has emerged as an alternative, but it still does not improve over the traditional method's performance. Here we introduce a framework to improve credit scoring models by blending several Graph Representation Learning methods: feature engineering, graph embeddings, and graph neural networks. We stacked their outputs to produce a single score in this approach. We validated this framework using a unique multi-source dataset that characterizes the relationships and credit history for the entire population of a Latin American country, applying it to credit risk models, application, and behavior, targeting both individuals and companies. Our results show that the graph representation learning methods should be used as complements, and these should not be seen as self-sufficient methods as is currently done. In terms of AUC and KS, we enhance the statistical performance, outperforming traditional methods. In Corporate lending, where the gain is much higher, it confirms that evaluating an unbanked company cannot solely consider its features. The business ecosystem where these firms interact with their owners, suppliers, customers, and other companies provides novel knowledge that enables financial institutions to enhance their creditworthiness assessment. Our results let us know when and which group to use graph data and what effects on performance to expect. They also show the enormous value of graph data on the unbanked credit scoring problem, principally to help companies' banking.
연구 동기 및 목표
- 신용 기록이 부족한 개인 및 기업(이른바 '얇은 파일' 대상)의 신용 평가 과제를 해결한다.
- 다양한 기법을 융합함으로써 고립된 그래프 표현 학습(GRL) 방법의 한계를 극복한다.
- 다양한 소스의 전국적 네트워크 및 금융 데이터를 활용해 개인 및 기업의 신용 평가 예측 성능을 향상시킨다.
- 사회적 네트워크 데이터가 신용도 평가에 가장 큰 기여를 하는 시점과 장소를 규명한다.
- 해석 가능한 인공지능(SHAP)을 활용해 특징 중요도 및 모델 행동에 대한 실질적인 통찰을 제공한다.
제안 방법
- 수작업 특징, Node2Vec 임베딩, 그래프 신경망(GNNs)을 통합한 세 가지 GRL 방법을 적용한 통합 정보 처리 프레임워크를 제안한다.
- 각 GRL 방법의 예측 결과를 하나의 입력 벡터로 통합하여 기울기 부스팅 분류기(XGBoost)에 입력한다.
- 라틴 아메리카 한 국가의 전체 인구를 포함하는 고유한 전국적 데이터셋을 사용하며, 금융 거래, 사회적 상호작용, 기업 관계를 포함한다.
- 개인 및 기업의 신용 평가, 행동 평가 등 네 가지 신용 평가 시나리오에 프레임워크를 적용한다.
- SHAP 값으로 특징 기여도 및 모델 결정 원리를 해석하여, 어떤 네트워크 특징이 예측에 가장 큰 영향을 미치는지 파악한다.
- 표준 평가 지표를 사용해 성능을 검증한다: 수신기 작동 특성 곡선 아래 면적(AUC) 및 코모고로프-스미르노프 통계량(KS)
실험 결과
연구 질문
- RQ1다양한 그래프 표현 학습(GRL) 기법을 융합할 경우, 각각을 별도로 사용할 때보다 신용 평가 성능이 향상되는가?
- RQ2통합된 네트워크 특징은 신용 리스크에 대해 어떤 통찰을 제공하며, 의사결정 과정을 어떻게 향상시키는가?
- RQ3개인 또는 기업의 신용 평가에서 사회적 네트워크 데이터가 가장 큰 성능 향상을 가져오는 분야는 어디인가?
- RQ4예를 들어, 에고 네트워크, 가족 관계, 공급망 등 어떤 종류의 네트워크 특징이 신용도 예측에 가장 유용한가?
- RQ5다른 GRL 방법(수작업 특징 공학, Node2Vec, GNNs)의 기여도는 대출자 유형 및 평가 시나리오에 따라 어떻게 달라지는가?
주요 결과
- 수작업 특징과 GNNs의 조합이 가장 높은 성능 향상을 이끌었으며, AUC 및 KS 지표에서 기준 모델을 크게 앞서갔다.
- 은행 서비스 미사용자 대상의 신용 평가에서 제안된 프레임워크가 가장 큰 성능 향상을 보였으며, 금융적으로 배제된 집단에 대한 가치를 입증했다.
- 기업 대출 분야에서는 기업의 소유주, 공급업체, 고객 등 기업 생태계가 기업 자체의 특성 외에도 중요한 예측 신호를 제공한다.
- BenchScore 기준 모델은 여전히 중요하지만, 네트워크 특징과 융합될 경우 상대적 기여도가 감소하여, 네트워크 데이터가 기존 지표를 대체하기보다는 보완한다는 점을 시사한다.
- Node2Vec 임베딩은 모델 성능 향상에 거의 기여하지 않아, 이 맥락에서 단독으로 특징 공학 방법으로서의 유용성이 제한됨을 시사한다.
- 개인 신용 평가에서는 가족 네트워크 특징(FamilyNet)이 더 넓은 사회적 네트워크(EOWNet)보다 더 유용한 정보를 제공함을 확인하여, 개인의 신용 평가에서 친족 관계의 중요성을 입증했다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.