Skip to main content
QUICK REVIEW

[论文解读] Integrating Stock Features and Global Information via Large Language Models for Enhanced Stock Return Prediction

Yujie Ding, Shuai Jia|arXiv (Cornell University)|Oct 9, 2023
Stock Market Forecasting MethodsDecision Sciences被引用 3
一句话总结

本文提出SCRL-LG框架,通过大型语言模型(LLMs)整合股票特征与全局信息,以提升股票收益预测性能。该方法采用局部-全局模型分解收益,并利用自相关强化学习将LLM生成的新闻嵌入与量化股票特征对齐于共享语义空间,从而在排名信息系数和异常收益方面实现卓越表现,尤其在A股市场较长持有期内表现更优。

ABSTRACT

The remarkable achievements and rapid advancements of Large Language Models (LLMs) such as ChatGPT and GPT-4 have showcased their immense potential in quantitative investment. Traders can effectively leverage these LLMs to analyze financial news and predict stock returns accurately. However, integrating LLMs into existing quantitative models presents two primary challenges: the insufficient utilization of semantic information embedded within LLMs and the difficulties in aligning the latent information within LLMs with pre-existing quantitative stock features. We propose a novel framework consisting of two components to surmount these challenges. The first component, the Local-Global (LG) model, introduces three distinct strategies for modeling global information. These approaches are grounded respectively on stock features, the capabilities of LLMs, and a hybrid method combining the two paradigms. The second component, Self-Correlated Reinforcement Learning (SCRL), focuses on aligning the embeddings of financial news generated by LLMs with stock features within the same semantic space. By implementing our framework, we have demonstrated superior performance in Rank Information Coefficient and returns, particularly compared to models relying only on stock features in the China A-share market.

研究动机与目标

  • 解决LLMs在股票收益预测中语义知识利用不足的问题。
  • 解决LLM生成的新闻嵌入与量化股票特征在不同语义空间中的对齐偏差问题。
  • 通过统一框架整合局部(个股)与全局(市场、行业、政策)信息,提升预测性能。
  • 证明LLMs在多模态股票收益预测中的有效性,超越情感分析的局限。

提出的方法

  • 局部-全局(LG)模型将股票收益分解为局部分量(α),由成交量和价格特征决定,以及全局分量(β),由市场、行业和政策信息决定。
  • 提出三种全局建模策略:(1)基于股票特征,(2)LLM原生,(3)混合,实现全局上下文的灵活整合。
  • 自相关强化学习(SCRL)通过对比学习目标,将LLM生成的新闻嵌入与股票特征对齐于共享语义空间。
  • 框架采用联合损失函数,结合排名损失与对比损失,以优化对齐效果与预测准确性。
  • 基于2022年中国A股数据进行回测,评估不同持有期下的性能表现。
  • 模型不仅用于情感分析,更用于从新闻标题中提取丰富的语义表征。

实验结果

研究问题

  • RQ1LLMs能否被有效用于情感分析之外的全面语义知识提取,以提升股票收益预测性能?
  • RQ2如何将LLM生成的嵌入与量化股票特征对齐于共享语义空间,以提升预测性能?
  • RQ3通过LLMs整合全局信息(市场、行业、政策)是否能缓解长期收益预测中的性能下降问题?
  • RQ4与仅依赖股票特征的模型相比,SCRL-LG框架在排名信息系数与异常收益方面表现如何?

主要发现

  • SCRL-LG模型在中国A股数据上,相较于仅使用股票特征的模型,在排名信息系数与异常收益方面均显著更优。
  • 随着预测周期延长,该模型保持或提升性能,而仅使用局部模型的版本则在年化收益率与夏普比率上显著下降。
  • LG-LLM与LG-STOCK变体在长持有期内表现出强鲁棒性,性能下降极小。
  • SCRL-LG模型在性能提升的同时,换手率持续下降,表明其交易信号更加高效且稳定。
  • 混合全局建模策略(结合股票特征与LLMs)在捕捉多层级(宏观、中观、微观)影响方面最为有效。
  • 该框架证实,将LLM生成的嵌入与股票特征对齐于统一语义空间,可显著提升预测准确性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。