[论文解读] Towards a Psychology of Machines: Large Language Models Predict Human Memory
本研究表明,大型语言模型(LLMs),如ChatGPT,能够准确预测在情境依赖记忆任务中的人类记忆表现。通过评估带有合适或不合适语境的拐弯句(garden-path sentences)的相关性与可记忆性,LLMs的评分与人类评分高度一致,并成功预测了后续的记忆回忆,表明LLMs能够模拟人类认知过程,为机器心理学这一新领域铺平道路。
Large language models (LLMs), such as ChatGPT, have shown remarkable abilities in natural language processing, opening new avenues in psychological research. This study explores whether LLMs can predict human memory performance in tasks involving garden-path sentences and contextual information. In the first part, we used ChatGPT to rate the relatedness and memorability of garden-path sentences preceded by either fitting or unfitting contexts. In the second part, human participants read the same sentences, rated their relatedness, and completed a surprise memory test. The results demonstrated that ChatGPT's relatedness ratings closely matched those of the human participants, and its memorability ratings effectively predicted human memory performance. Both LLM and human data revealed that higher relatedness in the unfitting context condition was associated with better memory performance, aligning with probabilistic frameworks of context-dependent learning. These findings suggest that LLMs, despite lacking human-like memory mechanisms, can model aspects of human cognition and serve as valuable tools in psychological research. We propose the field of machine psychology to explore this interplay between human cognition and artificial intelligence, offering a bidirectional approach where LLMs can both benefit from and contribute to our understanding of human cognitive processes.
研究动机与目标
- 探究大型语言模型(LLMs)是否能够预测在语境复杂句子任务中的人类记忆表现。
- 检验LLMs对相关性与可记忆性的评分是否与人类判断一致。
- 评估LLMs生成的可记忆性评分是否能预测实际的人类记忆回忆表现。
- 探索LLMs作为心理学研究工具的潜力,通过建模人类认知的某些方面。
- 基于人类认知与人工智能之间的双向互动,提出一个新跨学科领域——机器心理学。
提出的方法
- 通过提示让LLMs(特别是ChatGPT)对先前配有合适或不合适语境的拐弯句进行相关性与可记忆性评分。
- 人类被试阅读相同的句子,对相关性进行评分,并完成一次意外记忆测试以评估回忆表现。
- 使用相关性与回归分析比较LLMs与人类的相关性与可记忆性评分。
- 研究采用被试内设计,使用配对的句子-语境对,以确保有效比较。
- 统计建模评估LLMs生成的可记忆性评分是否能预测人类记忆表现。
- 采用情境依赖学习的概率框架,解释语境相关性与记忆结果之间的关系。
实验结果
研究问题
- RQ1大型语言模型能否准确预测具有语境差异的句子的人类记忆表现?
- RQ2LLMs生成的相关性与可记忆性评分与人类被试评分的接近程度如何?
- RQ3LLMs生成的可记忆性评分在多大程度上能预测实际的人类回忆表现?
- RQ4不合适的语境的相关性是否能提升记忆表现,LLMs能否检测到这一效应?
- RQ5LLMs能否作为心理学实验中人类认知过程的可靠替代指标?
主要发现
- LLMs对相关性的评分与人类被试评分高度相关,表明在感知与认知判断上具有强烈一致性。
- LLMs生成的可记忆性评分显著预测了人类在意外回忆测试中的记忆表现。
- 语境相关性越高——尤其是在不合适的语境条件下——记忆表现越好,这一模式被LLMs与人类共同捕捉到。
- 研究结果支持情境依赖学习的概率模型,即语境连贯性可增强记忆保持。
- 尽管缺乏生物记忆机制,LLMs仍表现出建模类人认知效应的能力。
- 研究结果表明,LLMs可作为有效且可扩展的工具,用于预测心理学研究中的人类认知反应。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。