[论文解读] AI-Augmented Predictions: LLM Assistants Improve Human Forecasting Accuracy
本研究评估了大型语言模型(LLMs)作为人类预测中的决策辅助工具,表明无论是高质量的'超级预测'助手,还是具有偏见、过度自信的助手,与使用较旧的非预测性LLM的对照组相比,都能平均将人类预测准确性提高23%。即使LLM存在缺陷,这种提升效果依然存在,表明LLM增强能提升推理能力,无论模型的可靠性如何。
Large language models (LLMs) match and sometimes exceeding human performance in many domains. This study explores the potential of LLMs to augment human judgement in a forecasting task. We evaluate the effect on human forecasters of two LLM assistants: one designed to provide high-quality ("superforecasting") advice, and the other designed to be overconfident and base-rate neglecting, thus providing noisy forecasting advice. We compare participants using these assistants to a control group that received a less advanced model that did not provide numerical predictions or engaged in explicit discussion of predictions. Participants (N = 991) answered a set of six forecasting questions and had the option to consult their assigned LLM assistant throughout. Our preregistered analyses show that interacting with each of our frontier LLM assistants significantly enhances prediction accuracy by between 24 percent and 28 percent compared to the control group. Exploratory analyses showed a pronounced outlier effect in one forecasting item, without which we find that the superforecasting assistant increased accuracy by 41 percent, compared with 29 percent for the noisy assistant. We further examine whether LLM forecasting augmentation disproportionately benefits less skilled forecasters, degrades the wisdom-of-the-crowd by reducing prediction diversity, or varies in effectiveness with question difficulty. Our data do not consistently support these hypotheses. Our results suggest that access to a frontier LLM assistant, even a noisy one, can be a helpful decision aid in cognitively demanding tasks compared to a less powerful model that does not provide specific forecasting advice. However, the effects of outliers suggest that further research into the robustness of this pattern is needed.
研究动机与目标
- 评估LLMs是否能改善现实世界中的前瞻性预测任务中的人类预测准确性。
- 检验LLM增强是否对技能较低的预测者带来更大益处。
- 调查LLM增强是否会降低预测多样性或损害聚合预测中的群体智慧。
- 评估LLM增强的效果是否随问题难度而变化。
- 确定LLM辅助带来的益处主要源于模型本身的准确性,还是源于认知增强机制。
提出的方法
- 参与者(N = 991)被随机分配至三种条件之一:使用GPT-3.5-turbo(DaVinci-003)的对照组、'超级预测'LLM助手,或具有偏见、过度自信的LLM助手。
- 所有参与者完成了涉及经济与市场指标的系列前瞻性预测任务,例如通货膨胀里程碑和石油储量。
- LLM助手被提示提供高质量、校准准确的预测,或表现出过度自信和基率忽视。
- 预测准确性通过Brier分数衡量,准确性越高,Brier分数越低。
- 预先注册的分析比较了各组之间的准确性,而探索性分析则考察了异常值和子群效应。
- 本研究采用受控实验设计,包含随机分组和盲态数据分析,以确保有效性。
实验结果
研究问题
- RQ1与使用能力较弱LLM的对照组相比,LLM增强是否显著提高人类预测准确性?
- RQ2技能较低的预测者是否从LLM增强中获得比高技能预测者更大的益处?
- RQ3LLM增强是否会降低预测多样性或损害聚合预测的准确性?
- RQ4LLM增强的效果是否随预测问题的难度而变化?
- RQ5预测准确性提升是源于LLM本身的预测质量,还是源于认知增强效应?
主要发现
- 与使用能力较弱LLM的对照组相比,LLM增强使人类预测准确性平均提高23%。
- 在排除一个异常值问题的探索性分析中,'超级预测'LLM助手使准确性提高43%,而具有偏见的助手使准确性提高28%。
- 尽管存在缺陷,具有偏见的LLM助手仍提高了预测准确性,表明即使模型不完美,也能提供有意义的认知增强。
- 低技能与高技能预测者在LLM增强中的获益无显著差异,挑战了技能较低者获益更大的假设。
- LLM增强未显著降低预测多样性,也未损害聚合预测的准确性,表明对集体智慧无一致的负面影响。
- LLM增强在简单与困难预测问题中的效果无显著差异,表明其益处均匀分布于不同任务难度。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。