[論文レビュー] Centaur: a foundation model of human cognition
CentaurはPsych-101で大規模言語モデルをファインチューニングして得られた人間認知の基盤モデルであり、幅広い実験を通じて人間の行動を予測・模倣するだけでなく神経データとも整合する。
Establishing a unified theory of cognition has been a major goal of psychology. While there have been previous attempts to instantiate such theories by building computational models, we currently do not have one model that captures the human mind in its entirety. A first step in this direction is to create a model that can predict human behavior in a wide range of settings. Here we introduce Centaur, a computational model that can predict and simulate human behavior in any experiment expressible in natural language. We derived Centaur by finetuning a state-of-the-art language model on a novel, large-scale data set called Psych-101. Psych-101 reaches an unprecedented scale, covering trial-by-trial data from over 60,000 participants performing over 10,000,000 choices in 160 experiments. Centaur not only captures the behavior of held-out participants better than existing cognitive models, but also generalizes to new cover stories, structural task modifications, and entirely new domains. Furthermore, we find that the model's internal representations become more aligned with human neural activity after finetuning. Taken together, our results demonstrate that it is possible to discover computational models that capture human behavior across a wide range of domains. We believe that such models provide tremendous potential for guiding the development of cognitive theories and present a case study to demonstrate this.
研究の動機と目的
- 統一的で汎ドメインの人間認知モデルの追求に動機づける。
- Psych-101を試行ごとの行動データセットとして大規模に紹介する。
- Centaurが多くの実験でドメイン特化モデルより人間の行動を予測できることを示す。
- Centaurが新しいcover story、タスク構造、および領域に一般化できることを示す。
- Centaurの内部表現が人間の神経活動と整合するかを検討する。
提案手法
- 非埋め込み層のアダプターを用いた量子化低ランク適応(QLoRA)で Psych-101 上で最先端の言語モデル (Llama 3.1 70B) をファインチューニングする。
- 160の心理学実験を自然言語プロンプトに書き起こして、試行ごとの履歴をカバーする Psych-101 を準備する。
- 交差エントロピー損失で1エポック訓練し、人間以外の応答トークンをマスキング、A100 GPU上で約5日。
- Hold-outな参加者と実験を横断してCentaurとLlamaおよびドメイン特化の認知モデルを比較するための疑似R^2 指標で評価する。
- Centaurが人間のような軌道を生成するかを評価するオープンループシミュレーションを実施する。
- 修正されたストーリー、タスク構造、および新規ドメインでの分布外テストを通じて一般化を評価する。
- Centaurの内部表現からfMRI信号を予測して人間データと比較することで神経整合性を分析する。

実験結果
リサーチクエスチョン
- RQ1Centaurは多様な実験において、ドメイン特化の認知モデルよりも保持データ外の人間の行動を予測できるか?
- RQ2Centaurは未見の実験、カバー・ストーリー、タスク構造を含む完全に新しいドメインに一般化するか?
- RQ3Finetuning後、Centaurの内部表現は人間の神経活動とより整合するようになるか?
- RQ4Centaurのオープンループシミュレーションは人間の軌道分布と一致するか?
- RQ5分布外評価におけるCentaurの評価は基準モデルとどうか?
主な発見
- Centaurはほとんどの実験でベースモデルのLlamaおよびドメイン特化の認知モデルの両方を上回る。
- ファインチューニングはLlamaに比べ平均で約0.14、ドメイン特化モデルに比べ約0.18のpseudo-R^2の改善をもたらした。
- Centaurのオープンループシミュレーションは人間のような軌道分布を生成し、モデルベース、モデルフリー、混合強化学習パターンを含む。
- Centaurは修正されたカバー・ストーリー、追加のタスク構造、および全く新しいドメインへ頑健に一般化して高い性能を示す。
- Centaurの内部表現は人間の神経活動とより整合し、タスク間でfMRI分析におけるデコーダビリティを向上させる。
![Figure 3: Evaluation in different held-out settings. a , Pseudo-R 2 values for the two-step task with a modified cover story [ 24 ] . b , Pseudo-R 2 values for a three-armed bandit experiment [ 25 ] . c , Pseudo-R 2 values for an experiment probing logical reasoning [ 26 ] . Centaur outperforms both](https://ar5iv.labs.arxiv.org/html/2410.20268/assets/x2.png)
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。