Skip to main content
QUICK REVIEW

[论文解读] BizFinBench.v2: A Unified Dual-Mode Bilingual Benchmark for Expert-Level Financial Capability Alignment

Xin Guo, Rongjunchen Zhang|arXiv (Cornell University)|Jan 10, 2026
Stock Market Forecasting Methods被引用 0
一句话总结

BizFinBench.v2 是一个大规模双语基准,使用真实的中国与美国市场数据,包含离线与在线任务,以评估大模型在现实世界金融场景中的能力。

ABSTRACT

Large language models have undergone rapid evolution, emerging as a pivotal technology for intelligence in financial operations. However, existing benchmarks are often constrained by pitfalls such as reliance on simulated or general-purpose samples and a focus on singular, offline static scenarios. Consequently, they fail to align with the requirements for authenticity and real-time responsiveness in financial services, leading to a significant discrepancy between benchmark performance and actual operational efficacy. To address this, we introduce BizFinBench.v2, the first large-scale evaluation benchmark grounded in authentic business data from both Chinese and U.S. equity markets, integrating online assessment. We performed clustering analysis on authentic user queries from financial platforms, resulting in eight fundamental tasks and two online tasks across four core business scenarios, totaling 29,578 expert-level Q&A pairs. Experimental results demonstrate that ChatGPT-5 achieves a prominent 61.5% accuracy in main tasks, though a substantial gap relative to financial experts persists; in online tasks, DeepSeek-R1 outperforms all other commercial LLMs. Error analysis further identifies the specific capability deficiencies of existing models within practical financial business contexts. BizFinBench.v2 transcends the limitations of current benchmarks, achieving a business-level deconstruction of LLM financial capabilities and providing a precise basis for evaluating efficacy in the widespread deployment of LLMs within the financial domain. The data and code are available at https://github.com/HiThink-Research/BizFinBench.v2.

研究动机与目标

  • 通过真实世界的中美市场数据,捕捉真实的金融商业能力。
  • 弥合离线基准与在线、实时金融服务需求之间的差距。
  • 提供双轨评估框架:核心商业能力 + 在线表现。
  • 通过基于专家的错误分析,识别大模型的核心能力不足之处。

提出的方法

  • 从真实市场数据中构建四个核心商业场景的八项离线任务和两项在线任务。
  • 将任务分为 Core Business Capabilities 和 Online Performance 以实现双轨评估。
  • 应用严格的三层质量控制(平台聚类、前线评审、专家交叉验证)以确保数据质量和合规性。
  • 对 Stock Price Prediction 和 Portfolio Asset Allocation 任务使用实时在线数据。
  • 在零-shot 设置下评估 21 个大模型(专有和开源),并对 SA 与 SPP 任务采用一致性预测。
  • 提供一个开源的大模型投资系统,以实现在线评估的可复现性。
Figure 1: BizFinBench.v2 comprises eight foundational tasks and two online tasks distributed across four major scenarios. The top-right corner displays a real-time screenshot of the Portfolio Asset Allocation task.
Figure 1: BizFinBench.v2 comprises eight foundational tasks and two online tasks distributed across four major scenarios. The top-right corner displays a real-time screenshot of the Portfolio Asset Allocation task.

实验结果

研究问题

  • RQ1离线情境中,基于中美市场的真实金融任务,LLMs 的表现如何?
  • RQ2在在线、实时金融任务(如股票价格预测与资产配置)中,LLMs 的能力如何?
  • RQ3在实际金融业务情境中,LLMs 的常见错误模式是什么,如何缓解?
  • RQ4在真实世界金融数据上,通用型与金融专业化 LLM 的表现有何差异?

主要发现

  • ChatGPT-5 在主要离线任务中达到最高的平均准确率(61.5%)。
  • 在在线任务中,DeepSeek-R1 超越所有其他商用大模型。
  • 开源模型 Qwen3-235B-A22B-Thinking-2507 在开源结果中以 53.3% 的平均准确率领先。
  • 金融专家在基础任务上的基线能力为 84.8%,高于当前大模型。
  • 错误分析揭示五大业务困境:Financial Semantic Deviation、Long-term Business Logic Discontinuity、MIAD、High-precision Computational Distortion、Financial Time-Series Logical Disorder。
  • 在资产配置指标(如总回报、夏普比率)方面,DeepSeek-R1 在商用模型中表现突出;部分顶尖模型在该任务中仍难以超越 SPY 基准。
Figure 2: We have ranked the performance of the LLMs participating in the evaluation under the zero-shot setting, and these results reflect their authentic practical business capabilities.
Figure 2: We have ranked the performance of the LLMs participating in the evaluation under the zero-shot setting, and these results reflect their authentic practical business capabilities.

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。