Skip to main content
QUICK REVIEW

[论文解读] Centaur: a foundation model of human cognition

Marcel Binz, Elif Akata|arXiv (Cornell University)|Oct 26, 2024
Cognitive Science and Education Research被引用 5
一句话总结

Centaur 是一个基于 Psych-101,对大型语言模型进行微调得到的人类认知基础模型,能够预测和模拟广泛实验中的人类行为,甚至与神经数据对齐。

ABSTRACT

Establishing a unified theory of cognition has been a major goal of psychology. While there have been previous attempts to instantiate such theories by building computational models, we currently do not have one model that captures the human mind in its entirety. A first step in this direction is to create a model that can predict human behavior in a wide range of settings. Here we introduce Centaur, a computational model that can predict and simulate human behavior in any experiment expressible in natural language. We derived Centaur by finetuning a state-of-the-art language model on a novel, large-scale data set called Psych-101. Psych-101 reaches an unprecedented scale, covering trial-by-trial data from over 60,000 participants performing over 10,000,000 choices in 160 experiments. Centaur not only captures the behavior of held-out participants better than existing cognitive models, but also generalizes to new cover stories, structural task modifications, and entirely new domains. Furthermore, we find that the model's internal representations become more aligned with human neural activity after finetuning. Taken together, our results demonstrate that it is possible to discover computational models that capture human behavior across a wide range of domains. We believe that such models provide tremendous potential for guiding the development of cognitive theories and present a case study to demonstrate this.

研究动机与目标

  • 推动追求一个统一的、领域无关的人类认知模型。
  • 将 Psych-101 介绍为一个大规模、逐试次的行为数据集。
  • 证明 Centaur 在大量实验中对人类行为的预测优于领域特定模型。
  • 展示 Centaur 对新的故事背景、任务结构和领域具有泛化能力。
  • 研究 Centaur 的内部表征是否与人类神经活动对齐。

提出的方法

  • 在 Psych-101 上对最先进的语言模型(Llama 3.1 70B)进行微调,使用量化低秩适配(QLoRA),在非嵌入层使用适配器。
  • 通过将 160 个心理学实验转写成覆盖逐试历史的自然语言提示来准备 Psych-101。
  • 在 A100 GPU 上以交叉熵损失训练一个时期,屏蔽非人类响应的标记,大约持续 ~5 天。
  • 使用伪 R^2 指标评估,将 Centaur 与 Llama 和领域特定认知模型在保留的参与者和实验上进行比较。
  • 进行开环仿真以评估 Centaur 是否产生类似人类的轨迹。
  • 通过对修改后的故事、任务结构和新领域的分布外测试来评估泛化。
  • 通过用 Centaur 的内部表征预测 fMRI 信号并与人类数据比较来分析神经对齐。
Figure 2: Performance on Psych-101. a , Pseudo-R 2 values for different models across experiments. A value of zero corresponds to prediction at chance level while a value of one corresponds to perfect predictability of human responses. Missing bars indicate performance below chance level. Centaur ou
Figure 2: Performance on Psych-101. a , Pseudo-R 2 values for different models across experiments. A value of zero corresponds to prediction at chance level while a value of one corresponds to perfect predictability of human responses. Missing bars indicate performance below chance level. Centaur ou

实验结果

研究问题

  • RQ1在广泛的实验中,Centaur 能否比领域特定认知模型更好地预测被保留的人类行为?
  • RQ2Centaur 是否能泛化到未见过的实验、故事背景和任务结构,甚至包括全新的领域?
  • RQ3微调后,Centaur 的内部表征是否与人类神经活动更加对齐?
  • RQ4Centaur 行为的开环仿真是否与人类轨迹分布一致?
  • RQ5与基线模型相比,Centaur 在分布外评估中的表现如何?

主要发现

  • 在大多数实验中,Centaur 超越了基线模型(Llama)和一系列领域特定认知模型。
  • 微调在平均上相对于 Llama 提升的伪 R^2 大约为 0.14,相对于领域特定模型约为 0.18。
  • Centaur 的开环仿真产生类似人类的轨迹分布,包括基于模型、无模型和混合强化学习模式。
  • Centaur 对修改后的故事背景、额外任务结构以及全新领域具有稳健的泛化能力。
  • Centaur 的内部表征与人类神经活动对齐程度更高,在跨任务的 fMRI 分析中提高了解码能力。
Figure 3: Evaluation in different held-out settings. a , Pseudo-R 2 values for the two-step task with a modified cover story [ 24 ] . b , Pseudo-R 2 values for a three-armed bandit experiment [ 25 ] . c , Pseudo-R 2 values for an experiment probing logical reasoning [ 26 ] . Centaur outperforms both
Figure 3: Evaluation in different held-out settings. a , Pseudo-R 2 values for the two-step task with a modified cover story [ 24 ] . b , Pseudo-R 2 values for a three-armed bandit experiment [ 25 ] . c , Pseudo-R 2 values for an experiment probing logical reasoning [ 26 ] . Centaur outperforms both

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。