[논문 리뷰] Empowering Many, Biasing a Few: Generalist Credit Scoring through Large Language Models
다양한 온라인 금융 작업 전반에 걸친 일반 목적 신용 점수 평가를 위한 지시 조정된 LLM 프레임워크 CALM을 소개하고, 9개 데이터셋 벤치마크와 편향 분석 및 공개 리소스에 중점을 둡니다.
In the financial industry, credit scoring is a fundamental element, shaping access to credit and determining the terms of loans for individuals and businesses alike. Traditional credit scoring methods, however, often grapple with challenges such as narrow knowledge scope and isolated evaluation of credit tasks. Our work posits that Large Language Models (LLMs) have great potential for credit scoring tasks, with strong generalization ability across multiple tasks. To systematically explore LLMs for credit scoring, we propose the first open-source comprehensive framework. We curate a novel benchmark covering 9 datasets with 14K samples, tailored for credit assessment and a critical examination of potential biases within LLMs, and the novel instruction tuning data with over 45k samples. We then propose the first Credit and Risk Assessment Large Language Model (CALM) by instruction tuning, tailored to the nuanced demands of various financial risk assessment tasks. We evaluate CALM, existing state-of-art (SOTA) methods, open source and closed source LLMs on the build benchmark. Our empirical results illuminate the capability of LLMs to not only match but surpass conventional models, pointing towards a future where credit scoring can be more inclusive, comprehensive, and unbiased. We contribute to the industry's transformation by sharing our pioneering instruction-tuning datasets, credit and risk assessment LLM, and benchmarks with the research community and the financial industry.
연구 동기 및 목표
- LLM이 단일 작업 전문가 시스템을 넘어서 다양한 온라인 신용 및 리스크 작업에 일반화할 수 있음을 증명한다.
- 신용 및 리스크 평가를 위한 9개 데이터셋(약 14K 샘플)의 포괄적 벤치마크를 만들고 공개한다.
- 대형 지시 조정 말뭉치를 활용하여 신용 및 리스크 작업에 맞춘 CALM 지시 조정 LLM을 개발한다.
- 신용 점수화 및 리스크 평가에 적용될 때 LLM의 잠재적 편향을 조사하고 윤리적 고려사항을 제시한다.
제안 방법
- 신용 점수, 사기 탐지, 재무 곤란, 청구 분석에 걸친 다양한 표형 데이터 벤치마크를 14K 샘플로 구성한다.
- 재샘플링으로 균형을 맞춘 45K 지시 조정 데이터셋(6개 데이터셋)을 구성하고 표 기반 및 설명 기반 형식의 프롬프트를 사용한다.
- LoRA를 사용한 5 에포크, AdamW, 학습률 3e-4, 가중치 감소 1e-5, 최대 입력 길이 2048로 LLaMa2-chat 모델을 미세조정한다.
- 정확도, F1, MCC, 편향 지표에서 CALM을 SOTA 전문가 시스템 및 다수의 오픈/비오픈 LLM과 비교 평가한다(예: GPT-4, ChatGPT, Bloomz, Vicuna, Llama 계열).
- AI FAIRNESS 360에 따라 데이터 편향(Disparate Impact)과 모델 편향(Equal Opportunity Difference, Average Odds Difference)을 분석한다.
실험 결과
연구 질문
- RQ1H1: 광범위한 사전 학습을 활용해 전통적인 신용/리스크 시스템의 좁은 전문성을 극복하고 다양한 온라인 작업에 일반화할 수 있는가?
- RQ2H2: 지시 조정된 LLM이 금융 데이터로 미세조정함으로써 여러 관련 신용 작업을 일반화/적응할 수 있는가?
- RQ3H3: LLM 능력의 발전이 신용 결정에서 공정성 편향을 도입하거나 확대하는가?
주요 결과
- LLMs, 특히 GPT-4는 여러 신용/리스크 작업에서 일부 기존 모델과 동등하거나 이를 능가할 수 있다.
- CALM은 미세조정된 LLM으로서 여러 신용/리스크 작업 간 지식 이전을 보여주고 훈련되지 않은 데이터셋에서 성능을 향상시킨다.
- 민감한 속성에서 LLM의 편향이 여전히 관찰되며, 배포 시 윤리적 감독의 필요성을 강조한다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.