Skip to main content
QUICK REVIEW

[Paper Review] AI-Augmented Predictions: LLM Assistants Improve Human Forecasting Accuracy

Philipp Schoenegger, Peter S. Park|arXiv (Cornell University)|Feb 12, 2024
Forecasting Techniques and ApplicationsDecision Sciences3 citations
TL;DR

This study evaluates large language models (LLMs) as decision aids in human forecasting, demonstrating that both a high-quality 'superforecasting' assistant and a biased, overconfident assistant significantly improve human forecasting accuracy by 23% on average compared to a control using an older, non-forecasting LLM. The improvement persists even when the LLM is flawed, suggesting LLM augmentation enhances reasoning regardless of model reliability.

ABSTRACT

Large language models (LLMs) match and sometimes exceeding human performance in many domains. This study explores the potential of LLMs to augment human judgement in a forecasting task. We evaluate the effect on human forecasters of two LLM assistants: one designed to provide high-quality ("superforecasting") advice, and the other designed to be overconfident and base-rate neglecting, thus providing noisy forecasting advice. We compare participants using these assistants to a control group that received a less advanced model that did not provide numerical predictions or engaged in explicit discussion of predictions. Participants (N = 991) answered a set of six forecasting questions and had the option to consult their assigned LLM assistant throughout. Our preregistered analyses show that interacting with each of our frontier LLM assistants significantly enhances prediction accuracy by between 24 percent and 28 percent compared to the control group. Exploratory analyses showed a pronounced outlier effect in one forecasting item, without which we find that the superforecasting assistant increased accuracy by 41 percent, compared with 29 percent for the noisy assistant. We further examine whether LLM forecasting augmentation disproportionately benefits less skilled forecasters, degrades the wisdom-of-the-crowd by reducing prediction diversity, or varies in effectiveness with question difficulty. Our data do not consistently support these hypotheses. Our results suggest that access to a frontier LLM assistant, even a noisy one, can be a helpful decision aid in cognitively demanding tasks compared to a less powerful model that does not provide specific forecasting advice. However, the effects of outliers suggest that further research into the robustness of this pattern is needed.

Motivation & Objective

  • To evaluate whether LLMs can improve human forecasting accuracy in real-world, prospective prediction tasks.
  • To examine whether LLM augmentation disproportionately benefits less skilled forecasters.
  • To investigate whether LLM augmentation reduces prediction diversity or degrades the wisdom of the crowd in aggregated forecasts.
  • To assess whether the effectiveness of LLM augmentation varies with question difficulty.
  • To determine whether the benefits of LLM assistance stem primarily from model accuracy or from cognitive augmentation mechanisms.

Proposed method

  • Participants (N = 991) were assigned to one of three conditions: a control group using GPT-3.5-turbo (DaVinci-003), a 'superforecasting' LLM assistant, or a biased, overconfident LLM assistant.
  • All participants completed a series of prospective forecasting tasks involving economic and market indicators, such as inflation milestones and oil reserves.
  • The LLM assistants were prompted to either provide high-quality, calibrated forecasts or to exhibit overconfidence and base-rate neglect.
  • Forecasting accuracy was measured using Brier scores, with higher accuracy corresponding to lower Brier scores.
  • Preregistered analyses compared accuracy across conditions, while exploratory analyses examined outliers and subgroup effects.
  • The study used a controlled experimental design with random assignment and blinded data analysis to ensure validity.

Experimental results

Research questions

  • RQ1Does LLM augmentation significantly improve human forecasting accuracy compared to a control using a less capable LLM?
  • RQ2Do less skilled forecasters benefit more from LLM augmentation than high-skilled forecasters?
  • RQ3Does LLM augmentation reduce prediction diversity or degrade the accuracy of aggregated forecasts?
  • RQ4Does the effectiveness of LLM augmentation vary with the difficulty of the forecasting question?
  • RQ5Is the improvement in forecasting accuracy due to the LLM’s inherent prediction quality or to cognitive augmentation effects?

Key findings

  • LLM augmentation improved human forecasting accuracy by 23% on average compared to the control group using a less capable LLM.
  • The superforecasting LLM assistant increased accuracy by 43% in exploratory analyses, excluding an outlier question, while the biased assistant improved accuracy by 28%.
  • The biased LLM assistant still improved forecasting accuracy despite its flaws, indicating that even flawed LLMs can provide meaningful cognitive augmentation.
  • There was no statistically significant difference in the benefit of LLM augmentation between low-skilled and high-skilled forecasters, challenging the hypothesis of disproportionate benefit to less skilled individuals.
  • LLM augmentation did not significantly reduce prediction diversity or degrade the accuracy of aggregated forecasts, suggesting no consistent negative impact on collective intelligence.
  • The effectiveness of LLM augmentation did not significantly differ between easy and hard forecasting questions, indicating a uniform benefit across task difficulty.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.