Skip to main content
QUICK REVIEW

[论文解读] Large Language Model Prediction Capabilities: Evidence from a Real-World Forecasting Tournament

Philipp Schoenegger, Peter S. Park|arXiv (Cornell University)|Oct 17, 2023
Topic Modeling被引用 7
一句话总结

本研究通过将GPT-4纳入Metaculus平台的真实世界三个月预测竞赛,评估其预测能力,将其在843个二元问题上的概率预测结果与人类群体预测进行对比,涵盖多样化领域。结果显示,GPT-4的预测表现显著逊于人类群体的中位数预测水平,其准确性与50%基准线无显著差异,表明尽管架构先进,其在真实世界中的预测能力仍十分有限。

ABSTRACT

Accurately predicting the future would be an important milestone in the capabilities of artificial intelligence. However, research on the ability of large language models to provide probabilistic predictions about future events remains nascent. To empirically test this ability, we enrolled OpenAI's state-of-the-art large language model, GPT-4, in a three-month forecasting tournament hosted on the Metaculus platform. The tournament, running from July to October 2023, attracted 843 participants and covered diverse topics including Big Tech, U.S. politics, viral outbreaks, and the Ukraine conflict. Focusing on binary forecasts, we show that GPT-4's probabilistic forecasts are significantly less accurate than the median human-crowd forecasts. We find that GPT-4's forecasts did not significantly differ from the no-information forecasting strategy of assigning a 50% probability to every question. We explore a potential explanation, that GPT-4 might be predisposed to predict probabilities close to the midpoint of the scale, but our data do not support this hypothesis. Overall, we find that GPT-4 significantly underperforms in real-world predictive tasks compared to median human-crowd forecasts. A potential explanation for this underperformance is that in real-world forecasting tournaments, the true answers are genuinely unknown at the time of prediction; unlike in other benchmark tasks like professional exams or time series forecasting, where strong performance may at least partly be due to the answers being memorized from the training data. This makes real-world forecasting tournaments an ideal environment for testing the generalized reasoning and prediction capabilities of artificial intelligence going forward.

研究动机与目标

  • 评估GPT-4在真实世界情境下的预测能力,其中真实答案未知且未出现在训练数据中。
  • 检验GPT-4是否能在概率预测中超越或匹配人类群体预测的中位数准确率。
  • 探究GPT-4表现不佳的原因是否源于对中间概率的偏好,或源于在不确定性下推理的根本性局限。
  • 确立预测竞赛作为评估大语言模型通用推理与预测能力的稳健基准,独立于训练数据的记忆化现象。
  • 探讨GPT-4预测表现薄弱对人工智能安全的影响,特别是对高级AI系统长期规划与目标导向行为的启示。

提出的方法

  • 在为期三个月(2023年7月至10月)的时间内,向GPT-4提供提示,要求其对Metaculus预测平台上的843个二元问题生成概率预测。
  • 使用Brier评分作为标准指标评估概率预测的准确性,得分越低表示表现越好。
  • 通过预注册的统计检验比较GPT-4的平均Brier得分与同一竞赛中人类预测者的中位数Brier得分,检验两组均值是否相等。
  • 通过单样本t检验检验GPT-4的预测是否系统性地偏向50%概率,原假设为预测的平均值为50%。
  • 提出使用贝叶斯模型平均方法结合人类与大语言模型的预测,但本研究未实际实现该方法。
  • 本研究在开放科学框架(Open Science Framework)上预注册了分析计划与数据收集流程,以确保透明度与可复现性。

实验结果

研究问题

  • RQ1在真实世界竞赛环境中,GPT-4的概率预测表现是否显著不同于人类群体的中位数预测?
  • RQ2GPT-4的预测准确率是否明显劣于对所有问题均分配50%概率的无信息基准?
  • RQ3GPT-4是否表现出系统性地倾向于预测接近50%的中点概率,表明其无法表达强信心?
  • RQ4在多大程度上,真实世界的预测竞赛可作为评估大语言模型通用推理与预测能力的有效基准,且独立于记忆化现象?
  • RQ5GPT-4在预测任务中表现薄弱,对未来发展具备长期规划能力的AI系统有何启示?

主要发现

  • GPT-4的平均Brier得分显著高于人类群体预测的中位数Brier得分,表明其概率预测准确性更差。
  • GPT-4的预测准确率与对所有问题均分配50%概率的无信息基准无显著差异。
  • 数据不支持GPT-4系统性地偏向中点概率的假设,因其预测分布的均值在统计上不显著接近50%。
  • 在多个领域(包括大型科技、美国政治、病毒爆发及乌克兰冲突)中,GPT-4在预测表现上显著逊于人类群体的中位数水平。
  • 本研究证实,真实世界的预测竞赛是评估大语言模型泛化能力的有效且稳健的环境,因为真实答案未知且未出现在训练数据中。
  • 研究结果表明,当前的大语言模型(如GPT-4)缺乏可靠的预见能力,这对人工智能安全以及未来AI系统实现长期目标导向行为具有重要启示。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。