Skip to main content
QUICK REVIEW

[Paper Review] Towards a Psychology of Machines: Large Language Models Predict Human Memory

Markus Huff, Elanur Ulakçı|arXiv (Cornell University)|Mar 8, 2024
Topic Modeling4 citations
TL;DR

This study demonstrates that large language models (LLMs), such as ChatGPT, can accurately predict human memory performance in context-dependent memory tasks. By rating the relatedness and memorability of garden-path sentences with fitting or unfitting contexts, LLMs closely mirrored human ratings and successfully predicted subsequent memory recall, suggesting LLMs can model human cognitive processes and paving the way for a new field of machine psychology.

ABSTRACT

Large language models (LLMs), such as ChatGPT, have shown remarkable abilities in natural language processing, opening new avenues in psychological research. This study explores whether LLMs can predict human memory performance in tasks involving garden-path sentences and contextual information. In the first part, we used ChatGPT to rate the relatedness and memorability of garden-path sentences preceded by either fitting or unfitting contexts. In the second part, human participants read the same sentences, rated their relatedness, and completed a surprise memory test. The results demonstrated that ChatGPT's relatedness ratings closely matched those of the human participants, and its memorability ratings effectively predicted human memory performance. Both LLM and human data revealed that higher relatedness in the unfitting context condition was associated with better memory performance, aligning with probabilistic frameworks of context-dependent learning. These findings suggest that LLMs, despite lacking human-like memory mechanisms, can model aspects of human cognition and serve as valuable tools in psychological research. We propose the field of machine psychology to explore this interplay between human cognition and artificial intelligence, offering a bidirectional approach where LLMs can both benefit from and contribute to our understanding of human cognitive processes.

Motivation & Objective

  • To investigate whether large language models (LLMs) can predict human memory performance in contextually complex sentence tasks.
  • To examine if LLMs' ratings of relatedness and memorability align with human judgments.
  • To evaluate whether LLM-derived memorability scores predict actual human memory recall.
  • To explore the potential of LLMs as tools for psychological research by modeling aspects of human cognition.
  • To propose a new interdisciplinary field—machine psychology—based on bidirectional interaction between human cognition and artificial intelligence.

Proposed method

  • LLMs (specifically ChatGPT) were prompted to rate the relatedness and memorability of garden-path sentences preceded by either fitting or unfitting contextual sentences.
  • Human participants read the same sentences, rated relatedness, and completed a surprise memory test to assess recall performance.
  • Relatedness and memorability ratings from both LLMs and humans were compared using correlation and regression analyses.
  • The study employed a within-subjects design with matched sentence-context pairs to ensure valid comparisons.
  • Statistical modeling assessed whether LLM-generated memorability scores predicted human memory performance.
  • The probabilistic framework of context-dependent learning was used to interpret the relationship between context relatedness and memory outcomes.

Experimental results

Research questions

  • RQ1Can large language models accurately predict human memory performance for sentences with contextual variations?
  • RQ2How closely do LLM-generated ratings of relatedness and memorability align with human participant ratings?
  • RQ3To what extent do LLM-derived memorability scores predict actual human recall performance?
  • RQ4Does the relatedness of an unfitting context improve memory performance, and can LLMs detect this effect?
  • RQ5Can LLMs serve as reliable proxies for human cognitive processes in psychological experiments?

Key findings

  • LLM ratings of relatedness showed high correlation with human participant ratings, indicating strong alignment in perceptual and cognitive judgments.
  • LLM-generated memorability scores significantly predicted human memory performance in surprise recall tests.
  • Higher relatedness in the context—particularly in the unfitting condition—was associated with better memory performance, a pattern captured by both LLMs and humans.
  • The results support probabilistic models of context-dependent learning, where contextual coherence enhances memory retention.
  • LLMs demonstrated the ability to model human-like cognitive effects despite lacking biological memory mechanisms.
  • The findings suggest that LLMs can serve as valid, scalable tools for predicting human cognitive responses in psychological research.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.