[論文レビュー] AI-Augmented Predictions: LLM Assistants Improve Human Forecasting Accuracy
本研究では、大規模言語モデル(LLM)が人間の予測に与える意思決定支援の有効性を評価し、高品質な「スーパーフォーキャスティング」アシスタントおよび偏りがあり過信傾向の強いアシスタントの両方が、過去の非予測用LLMを使用する対照群と比較して、平均で23%の予測精度向上を示した。モデルの信頼性が低い場合でも改善効果が維持されることから、LLMの補完がモデルの信頼性にかかわらず推論を強化することを示唆している。
Large language models (LLMs) match and sometimes exceeding human performance in many domains. This study explores the potential of LLMs to augment human judgement in a forecasting task. We evaluate the effect on human forecasters of two LLM assistants: one designed to provide high-quality ("superforecasting") advice, and the other designed to be overconfident and base-rate neglecting, thus providing noisy forecasting advice. We compare participants using these assistants to a control group that received a less advanced model that did not provide numerical predictions or engaged in explicit discussion of predictions. Participants (N = 991) answered a set of six forecasting questions and had the option to consult their assigned LLM assistant throughout. Our preregistered analyses show that interacting with each of our frontier LLM assistants significantly enhances prediction accuracy by between 24 percent and 28 percent compared to the control group. Exploratory analyses showed a pronounced outlier effect in one forecasting item, without which we find that the superforecasting assistant increased accuracy by 41 percent, compared with 29 percent for the noisy assistant. We further examine whether LLM forecasting augmentation disproportionately benefits less skilled forecasters, degrades the wisdom-of-the-crowd by reducing prediction diversity, or varies in effectiveness with question difficulty. Our data do not consistently support these hypotheses. Our results suggest that access to a frontier LLM assistant, even a noisy one, can be a helpful decision aid in cognitively demanding tasks compared to a less powerful model that does not provide specific forecasting advice. However, the effects of outliers suggest that further research into the robustness of this pattern is needed.
研究の動機と目的
- 実世界の前向きな予測課題において、LLMが人間の予測精度を向上させることを評価すること。
- LLM補完が、スキルが低い予測者に比べてスキルが高い予測者よりも顕著に利益をもたらすかどうかを検討すること。
- LLM補完が予測の多様性を低下させたり、集約予測における「群衆の知恵」を損なうかどうかを調査すること。
- LLM補完の有効性が予測課題の難易度に応じて変化するかどうかを評価すること。
- LLM支援の恩恵が主にモデルの予測精度に起因するのか、それとも認知的補完メカニズムに起因するのかを特定すること。
提案手法
- 参加者(N = 991)は、3つの条件のいずれかに無作為に割り当てられた:GPT-3.5-turbo(DaVinci-003)を使用する対照群、『スーパーフォーキャスティング』LLMアシスタント、または偏りがあり過信傾向の強いLLMアシスタント。
- 全参加者がインフレーションのマイルストーンや石油埋蔵量といった経済的・市場指標を含む一連の前向き予測タスクを完了した。
- LLMアシスタントには、高品質でキャリブレーションされた予測を提供するか、過信とベースレート無視を示すようにプロンプトが与えられた。
- 予測精度はBrierスコアを用いて測定され、高い精度は低いBrierスコアに対応する。
- 事前に登録された分析では条件間の精度を比較したが、探索的分析では外れ値やサブグループ効果を検討した。
- 有効性を保証するため、無作為割り当てと盲検データ分析を含む制御された実験デザインが採用された。
実験結果
リサーチクエスチョン
- RQ1LLM補完は、能力が低いLLMを使用する対照群と比較して、人間の予測精度を顕著に向上させるか?
- RQ2スキルが低い予測者が、スキルが高い予測者よりもLLM補完の恩恵をより多く受けるか?
- RQ3LLM補完は予測の多様性を低下させたり、集約予測の精度を低下させるか?
- RQ4LLM補完の有効性は予測課題の難易度に応じて変化するか?
- RQ5予測精度の向上は、LLMの本質的予測品質に起因するのか、それとも認知的補完効果に起因するのか?
主な発見
- LLM補完は、能力が低いLLMを使用する対照群と比較して、平均で23%の予測精度向上を示した。
- 探索的分析では、外れ値を除いた結果、スーパーフォーキャスティングLLMアシスタントは精度を43%向上させたが、偏りのあるアシスタントは28%向上させた。
- 偏ったLLMアシスタントですら予測精度を向上させたことから、すなわち、欠陥のあるLLMでも意味のある認知的補完を提供できる可能性がある。
- スキルが低い予測者と高い予測者との間でLLM補完の恩恵に有意差は認められず、スキルが低い人への顕著な利益という仮説は支持されなかった。
- LLM補完は予測の多様性を顕著に低下させず、集約予測の精度を低下させないことも示され、集団知性に一貫した悪影響がないことが示唆された。
- 難易度が易しいか難しいかに関わらず、LLM補完の有効性に有意差は認められず、タスクの難易度にかかわらず一貫した利益が得られることを示した。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。