[논문 리뷰] BizFinBench.v2: A Unified Dual-Mode Bilingual Benchmark for Expert-Level Financial Capability Alignment
BizFinBench.v2는 실제 중국 및 미국 시장 데이터를 사용하고 오프라인 및 온라인 과제를 포함하는 대규모 양언어 벤치마크로, LLMs의 실세계 재무 역량을 평가합니다.
Large language models have undergone rapid evolution, emerging as a pivotal technology for intelligence in financial operations. However, existing benchmarks are often constrained by pitfalls such as reliance on simulated or general-purpose samples and a focus on singular, offline static scenarios. Consequently, they fail to align with the requirements for authenticity and real-time responsiveness in financial services, leading to a significant discrepancy between benchmark performance and actual operational efficacy. To address this, we introduce BizFinBench.v2, the first large-scale evaluation benchmark grounded in authentic business data from both Chinese and U.S. equity markets, integrating online assessment. We performed clustering analysis on authentic user queries from financial platforms, resulting in eight fundamental tasks and two online tasks across four core business scenarios, totaling 29,578 expert-level Q&A pairs. Experimental results demonstrate that ChatGPT-5 achieves a prominent 61.5% accuracy in main tasks, though a substantial gap relative to financial experts persists; in online tasks, DeepSeek-R1 outperforms all other commercial LLMs. Error analysis further identifies the specific capability deficiencies of existing models within practical financial business contexts. BizFinBench.v2 transcends the limitations of current benchmarks, achieving a business-level deconstruction of LLM financial capabilities and providing a precise basis for evaluating efficacy in the widespread deployment of LLMs within the financial domain. The data and code are available at https://github.com/HiThink-Research/BizFinBench.v2.
연구 동기 및 목표
- 실세계 중국 및 미국 시장 데이터를 활용하여 진정한 재무 비즈니스 역량을 포착합니다.
- 오프라인 벤치마크와 온라인 실시간 금융 서비스 수요 사이의 간극을 해소합니다.
- Core Business Capabilities + Online Performance라는 이중 트랙 평가 프레임워크를 제공합니다.
- 전문가 기반 오류 분석을 통해 LLMs의 핵심 역량 결함을 식별합니다.
제안 방법
- 실시장 데이터에서 네 가지 핵심 비즈니스 시나리오에 걸친 여덟 개의 오프라인 과제와 두 개의 온라인 과제를 구성합니다.
- 과제를 Core Business Capabilities와 Online Performance로 구성하여 이중 트랙 평가를 수행합니다.
- 데이터 품질과 규정을 보장하기 위해 엄격한 3단계 품질 관리(플랫폼 클러스터링, 최전선 검토, 전문가 교차 검증)를 적용합니다.
- 주가 예측 및 포트폴리오 자산 배분 과제에 실시간 online 데이터를 사용합니다.
- SA 및 SPP 과제에 대한 컨포멀 예측을 적용하여 제로샷 설정에서 21개의 LLMs(독점 및 오픈 소스)를 평가합니다.
- PAA에서 재현 가능한 온라인 평가를 위한 오픈 소스 LLM 투자 시스템을 제공합니다.

실험 결과
연구 질문
- RQ1오프라인 설정에서 중국 및 미국 시장에서 추출된 진짜 재무 과제에 대해 LLMs가 얼마나 잘 수행합니까?
- RQ2주가 예측 및 자산 배분과 같은 온라인 실시간 재무 과제에서 LLMs의 역량은 어느 정도입니까?
- RQ3실무 재무 비즈니스 맥락에서 LLMs의 일반적인 오류 모드는 무엇이며 어떻게 완화할 수 있습니까?
- RQ4실세계 재무 데이터에서 범용 LLMs와 금융 특화 LLMs 간의 성능 차이는 어떻게 나타납니까?
주요 결과
- ChatGPT-5가 주요 오프라인 과제에서 평균 정확도 61.5%로 가장 높습니다.
- 온라인 과제에서는 DeepSeek-R1이 모든 다른 상용 LLM을 능가합니다.
- 오픈 소스 모델 Qwen3-235B-A22B-Thinking-2507이 평균 정확도 53.3%로 오픈 소스 결과를 선도합니다.
- 재무 전문가들은 기초 과제에서 현재 LLMs보다 더 높은 기본 역량(84.8%)을 달성합니다.
- 오류 분석은 다섯 가지 비즈니스 딜레마를 드러냅니다: Financial Semantic Deviation, Long-term Business Logic Discontinuity, MIAD, High-precision Computational Distortion, and Financial Time-Series Logical Disorder.
- 상용 모델 중에서 자산 배분 지표(예: 총수익, Sharpe 비율)에서 DeepSeek-R1이 뛰어나며, 일부 상위 모델은 이 과제에서 SPY 벤치마크를 능가하는 데 어려움을 겪습니다.

더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.