[논문 리뷰] FinBen: A Holistic Financial Benchmark for Large Language Models
FinBen은 LLM을 평가하기 위해 35개 데이터셋과 23개의 금융 과제를 포함하는 포괄적인 오픈 소스 벤치마크를 도입하며, induction에서 trading까지의 세 가지 CHC에서 영감을 받은 스펙럼으로 구성되어 있다.
LLMs have transformed NLP and shown promise in various fields, yet their potential in finance is underexplored due to a lack of comprehensive evaluation benchmarks, the rapid development of LLMs, and the complexity of financial tasks. In this paper, we introduce FinBen, the first extensive open-source evaluation benchmark, including 36 datasets spanning 24 financial tasks, covering seven critical aspects: information extraction (IE), textual analysis, question answering (QA), text generation, risk management, forecasting, and decision-making. FinBen offers several key innovations: a broader range of tasks and datasets, the first evaluation of stock trading, novel agent and Retrieval-Augmented Generation (RAG) evaluation, and three novel open-source evaluation datasets for text summarization, question answering, and stock trading. Our evaluation of 15 representative LLMs, including GPT-4, ChatGPT, and the latest Gemini, reveals several key findings: While LLMs excel in IE and textual analysis, they struggle with advanced reasoning and complex tasks like text generation and forecasting. GPT-4 excels in IE and stock trading, while Gemini is better at text generation and forecasting. Instruction-tuned LLMs improve textual analysis but offer limited benefits for complex tasks such as QA. FinBen has been used to host the first financial LLMs shared task at the FinNLP-AgentScen workshop during IJCAI-2024, attracting 12 teams. Their novel solutions outperformed GPT-4, showcasing FinBen's potential to drive innovation in financial LLMs. All datasets, results, and codes are released for the research community: https://github.com/The-FinAI/PIXIU.
연구 동기 및 목표
- LLMs를 위한 광범위하고 실제 세계의 금융 평가 벤치마크가 필요하다는 점을 동기화한다.
- 언어 처리, 지식 추출, 수치 추론, 생성, 예측 및 거래 과제를 포괄하도록 FinBen을 설계한다.
- 금융에서 LLM의 능력을 평가하기 위해 다양한 데이터 모달리티를 갖춘 오픈 소스 프레임워크를 제공한다.
- 금융 과제에서 강점과 한계를 식별하기 위해 15개의 대표 LLM을 평가한다.
- 금융에서 기본에서 일반 지능까지의 인지 능력을 매핑하기 위한 CHC-inspired 스펙트럼을 제안한다.
제안 방법
- FinBen을 23개 금융 과제에 걸쳐 35개의 데이터셋으로 구성한다.
- CHC 이론을 모방하는 세 가지 스펙트럼으로 과제를 구성한다: Spectrum I (Quantification, Extraction, Numerical Understanding), Spectrum II (Generation, Forecasting), Spectrum III (Stock Trading).
- GPT-4, ChatGPT, Gemini를 포함한 15개 LLM의 zero-shot 및 few-shot 성능을 평가한다.
- 표준 지표를 과제별로 사용한다(예: F1, accuracy, RMSE, ROUGE/BERTScore/BARTScore, MCC, EMAcc) 및 거래 지표(CR, SR, DV, AV, MD).
- 지침 미세조정이 도움이 되는 영역과 여전히 간극이 존재하는 영역을 식별하기 위해 과제 간 성능을 비교한다.
실험 결과
연구 질문
- RQ1FinBen이 기존 NLP 중심 벤치마크를 넘어 금융 분야의 광범위하고 실제적인 LLM 평가를 제공할 수 있는가?
- RQ2현재 LLM이 어떤 금융 과제에서 우수하고, 어떤 영역에서 어려움을 겪는가(예: 복잡한 추출, 수치 추론, 예측)?
- RQ3다른 모델 계열(GPT-4, Gemini, 오픈 소스 LLM 포함)이 세 가지 CHC-inspired 스펙트럼에서 어떻게 비교되는가?
- RQ4지침 미세조정이 모든 과제에서 성능을 균일하게 향상시키는가, 아니면 더 간단한 과제에서만 효과적인가?
주요 결과
- GPT-4는 정량화, 추출, 수치적 추론 및 주식 거래에서 선두를 차지하고; Gemini은 생성 및 예측에서 우수하다.
- 지시 미세조정은 단순한 과제에서는 성능을 향상시키지만, 복잡한 수치 추론, 생성 및 예측에서는 효과가 덜하다.
- 오픈 소스/중국어 튜닝 모델은 일부 분류 과제에서 강력한 성능을 보이지만, 다언어 효과와 데이터셋 정렬이 결과에 영향을 준다.
- 주식 거래 과제에서 LLM의 일반 인지 능력을 드러내며, GPT-4가 평가된 모델 중 최고 샤프비율과 최소 최대 낙폭을 달성한다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.