[Paper Review] ChatGPT and Deepseek: Can They Predict the Stock Market and Macroeconomy?
ChatGPT can extract information from Wall Street Journal headlines to predict stock returns and the market risk premium, while DeepSeek underperforms; other LLMs also fall short.
We study whether ChatGPT and DeepSeek can extract information from the Wall Street Journal to predict the stock market and the macroeconomy. We find that ChatGPT has predictive power. DeepSeek underperforms ChatGPT, which is trained more extensively in English. Other large language models also underperform. Consistent with financial theories, the predictability is driven by investors' underreaction to positive news, especially during periods of economic downturn and high information uncertainty. Negative news correlates with returns but lacks predictive value. At present, ChatGPT appears to be the only model capable of capturing economic news that links to the market risk premium.
Motivation & Objective
- Assess whether ChatGPT and DeepSeek can forecast stock market returns and macroeconomic variables from Wall Street Journal front-page news headlines.
- Quantify the predictive power of good vs. bad news ratios derived from LLMs.
- Compare ChatGPT with DeepSeek and other large language models in forecasting performance.
- Examine economic mechanisms (investor underreaction, information uncertainty) driving any predictive power.
Proposed method
- Use Wall Street Journal front-page headlines (1996–2022) as input data.
- Prompt ChatGPT-3.5 to classify headlines as GOING UP, GOING DOWN, or UNKNOWN and compute monthly good/bad news ratios.
- Evaluate in-sample and out-of-sample predictability of market excess returns using the good-news ratio NR^G.
- Test robustness with ChatGPT-4, fine-tuning, and alternative prompts; compare with DeepSeek-R1 and BERT-family models.
- Analyze embedding-based novelty metrics and control for lagged factors and macro variables.
Experimental results
Research questions
- RQ1Can ChatGPT and DeepSeek extract information from WSJ headlines to predict the aggregate stock market and the market risk premium?
- RQ2Are good-news signals from ChatGPT predictive for future returns, and how do they perform out-of-sample?
- RQ3What are the comparative performances of ChatGPT, DeepSeek, and other LLMs in predicting stock returns and macroeconomic fundamentals?
Key findings
- ChatGPT-3.5’s good-news ratio NR^G positively predicts contemporaneous and future market returns, with R^2 rising to 8.52% over an annual horizon (Jan 1996–Dec 2022).
- Out-of-sample R_OS^2 for NR^G is 1.17% (Jan 2006–Dec 2022), with meaningful economic value (CER gain 4.92% at risk aversion=3; net-of-cost CER 3.55%; Sharpe ratio 0.51 vs 0.30 for the market).
- Negative news signals have contemporaneous correlation with returns but no predictive power for future returns; good news drives forecastability, especially in downturns and periods of high uncertainty.
- DeepSeek-R1 captures contemporaneous stock-market reactions to news but lacks forecasting power for future returns or macro fundamentals; its signals correlate with sentiment but differ from GPT-derived signals.
- ChatGPT outperforms other large language models (DeepSeek and BERT-family) in capturing macroeconomic information relevant to the market risk premium.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.