Skip to main content
QUICK REVIEW

[Paper Review] Can ChatGPT Forecast Stock Price Movements? Return Predictability and Large Language Models

Alejandro Lopez-Lira, Yuehua Tang|arXiv (Cornell University)|Apr 15, 2023
Stock Market Forecasting Methods8 citations
TL;DR

The paper evaluates whether ChatGPT and other LLMs can forecast stock returns from news headlines, finding a positive linkage and ChatGPT outperforming traditional sentiment methods, with ChatGPT-4 delivering the strongest results.

ABSTRACT

We document the capability of large language models (LLMs) like ChatGPT to predict stock market reactions from news headlines without direct financial training. Using post-knowledge-cutoff headlines, GPT-4 captures initial market responses, achieving approximately 90% portfolio-day hit rates for the non-tradable initial reaction. GPT-4 scores also significantly predict the subsequent drift, especially for small stocks and negative news. Forecasting ability generally increases with model size, suggesting that financial reasoning is an emerging capacity of complex LLMs. Strategy returns decline as LLM adoption rises, consistent with improved price efficiency. To rationalize these findings, we develop a theoretical model that incorporates LLM technology, information-processing capacity constraints, underreaction, and limits to arbitrage.

Motivation & Objective

  • Motivate the question of whether large language models can predict stock returns using textual information.
  • Assess ChatGPT and competing LLMs’ ability to extract signals from news headlines to forecast next-day returns.
  • Compare LLM-based signals to traditional vendor sentiment scores.
  • Quantify investment performance using long-short strategies and evaluate robustness to transaction costs.
  • Explore the capabilities of progressively more advanced models (GPT-1, GPT-2, BERT, ChatGPT variants) in return predictability.

Proposed method

  • Construct a dataset of US stock returns from CRSP and headlines from major news sources matched to RavenPack data.
  • Convert each headline into a ChatGPT score (YES=1, UNKNOWN=0, NO=-1) via a prescribed prompt and aggregate across headlines by day.
  • Run out-of-sample predictive regressions of next-day returns on ChatGPT scores and rival sentiment scores with firm and date fixed effects.
  • Form zero-cost long-short portfolios based on positive/negative ChatGPT signals and assess performance with and without transaction costs.
  • Compare performance across ChatGPT-3.5, ChatGPT-4, BART Large, and basic models (GPT-1, GPT-2, BERT).
  • Evaluate model reasoning using a new method that links recommendation correctness to its explicit reasoning words.

Experimental results

Research questions

  • RQ1Can ChatGPT-derived headline sentiment predict next-day stock returns beyond traditional sentiment measures?
  • RQ2Do more advanced LLMs (e.g., ChatGPT-4) yield stronger predictive power than earlier models and basic NLP models?
  • RQ3Is return predictability driven by market underreaction, with stronger effects for small caps and bad news?
  • RQ4Does incorporating LLM-based signals improve Sharpe ratios in practical trading strategies?

Key findings

  • ChatGPT-3.5 signaling is significantly related to next-day returns; a move from -1 to +1 predicts about 51.8 basis points in next-day returns.
  • A self-financing long-short strategy based on ChatGPT-3.5 yields cumulative returns over 550% from 2021-10 to 2022-12 without costs; with 10-25 bps transaction costs, cumulative returns are 350% and 50% respectively.
  • ChatGPT-4 long-short strategy delivers over 350% cumulative return with a Sharpe ratio of 3.8 and a max drawdown of -10.4%, outperforming ChatGPT-3.5 (Sharpe 3.1; drawdown -22.8%).
  • ChatGPT outperforms traditional vendor sentiment scores when both are included in regressions; vendor scores become insignificant.
  • Predictability is present for both small and large stocks, but is stronger for smaller stocks and for stocks with negative news, suggesting limits-to-arbitrage effects.
  • GPT-1, GPT-2, and BERT show little to no stock forecasting capability, highlighting the value added by larger, more capable models.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.