[论文解读] A Survey of Large Language Models in Finance (FinLLMs)
本综述追踪从通用领域的 LLMs 到金融领域的 FinPLMs 与 FinLLMs 的演变,比较技术,概述六个基准和八个高级任务,并讨论 FinLLMs 的机会与挑战。
Large Language Models (LLMs) have shown remarkable capabilities across a wide variety of Natural Language Processing (NLP) tasks and have attracted attention from multiple domains, including financial services. Despite the extensive research into general-domain LLMs, and their immense potential in finance, Financial LLM (FinLLM) research remains limited. This survey provides a comprehensive overview of FinLLMs, including their history, techniques, performance, and opportunities and challenges. Firstly, we present a chronological overview of general-domain Pre-trained Language Models (PLMs) through to current FinLLMs, including the GPT-series, selected open-source LLMs, and financial LMs. Secondly, we compare five techniques used across financial PLMs and FinLLMs, including training methods, training data, and fine-tuning methods. Thirdly, we summarize the performance evaluations of six benchmark tasks and datasets. In addition, we provide eight advanced financial NLP tasks and datasets for developing more sophisticated FinLLMs. Finally, we discuss the opportunities and the challenges facing FinLLMs, such as hallucination, privacy, and efficiency. To support AI research in finance, we compile a collection of accessible datasets and evaluation benchmarks on GitHub.
研究动机与目标
- 绘制从通用领域的 PLMs 到金融领域的 FinPLMs 与 FinLLMs 的历史演进。
- 比较 FinPLMs 与 FinLLMs 使用的训练与微调技术。
- 概述多个金融NLP任务和数据集上的基准性能。
- 介绍面向未来 FinLLM 发展的高级金融NLP任务与数据集。
- 讨论 FinLLMs 在真实金融应用中的机会、挑战与实际考虑因素。
提出的方法
- 综述从 GPT 系列与开源 LLMs 到 FinLLMs 和金融领域模型的演变。
- 比较四个 FinPLMs 和四个 FinLLMs 的五种技术,重点关注训练数据、方法和指示性微调。
- 概述六个基准任务和数据集上的性能,并勾勒八个高级金融NLP任务与数据集。
- 汇编可获取的 GitHub 数据集与基准,支持未来的 FinLLM 研究。
- 讨论 FinLLMs 的实际考虑因素,如隐私、效率和幻觉问题。

实验结果
研究问题
- RQ1从通用领域的 LM 到 FinLLMs 的历史进展是什么,哪些模型定义了这一轨迹?
- RQ2哪些训练、数据与微调技术表征 FinPLMs 与 FinLLMs?
- RQ3在既定的金融NLP 基准上 FinPLMs 与 FinLLMs 的表现如何,高级任务揭示了哪些差距?
- RQ4有哪些数据集和基准存在或需要用于推进 FinLLMs, 如何利用它们推进未来研究?
主要发现
- 跨领域的 FinPLMs 在情感分析、文本分类和命名实体识别(NER)任务上表现出色。
- 特定任务的 SOTA 模型在问答、SMP、和摘要等任务上超过 FinLLMs,表明在这些领域的 FinLLMs 还有提升空间。
- GPT-4 在大多数基准上表现强劲,除了摘要任务,在该任务上是任务特定模型占优。
- FinMA、InvestLM、FinGPT 与 BloombergGPT 展示了许可、数据源和架构选择(基于 LLaMA、BLOOM 风格等)的光谱。
- RAG 及其他基于检索的方法被强调为改进 FinLLMs 可靠性与隐私性的有前景方向。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。