[Paper Review] Event Segmentation Applications in Large Language Model Enabled Automated Recall Assessments
The paper demonstrates that large language models can automate event segmentation and recall assessments, aligning with human segmentation patterns and offering a scalable alternative to manual scoring.
Understanding how individuals perceive and recall information in their natural environments is critical to understanding potential failures in perception (e.g., sensory loss) and memory (e.g., dementia). Event segmentation, the process of identifying distinct events within dynamic environments, is central to how we perceive, encode, and recall experiences. This cognitive process not only influences moment-to-moment comprehension but also shapes event specific memory. Despite the importance of event segmentation and event memory, current research methodologies rely heavily on human judgements for assessing segmentation patterns and recall ability, which are subjective and time-consuming. A few approaches have been introduced to automate event segmentation and recall scoring, but validity with human responses and ease of implementation require further advancements. To address these concerns, we leverage Large Language Models (LLMs) to automate event segmentation and assess recall, employing chat completion and text-embedding models, respectively. We validated these models against human annotations and determined that LLMs can accurately identify event boundaries, and that human event segmentation is more consistent with LLMs than among humans themselves. Using this framework, we advanced an automated approach for recall assessments which revealed semantic similarity between segmented narrative events and participant recall can estimate recall performance. Our findings demonstrate that LLMs can effectively simulate human segmentation patterns and provide recall evaluations that are a scalable alternative to manual scoring. This research opens novel avenues for studying the intersection between perception, memory, and cognitive impairment using methodologies driven by artificial intelligence.
Motivation & Objective
- Motivate study of how people perceive and recall events in natural environments and its relevance to perception and memory impairments.
- Develop an automated pipeline for event segmentation using LLMs and assess recall using semantic similarity to segmented events.
- Validate LLM-based segmentation against human annotations and compare consistency between humans and LLMs.
- Show that automated recall scoring via LLMs can scale cognitive assessment in memory-related research.
Proposed method
- Use chat-based completion models to perform automated event segmentation from narratives.
- Use text embedding models to quantify recall similarity to segmented events.
- Validate automated segmentation against human annotations to assess validity.
- Compare consistency of human segmentation with LLM segmentation versus human-human consistency.
- Provide a framework where semantic similarity between segmented events and recall estimates serves as a recall measure.
Experimental results
Research questions
- RQ1Can LLMs accurately identify event boundaries in naturalistic narratives?
- RQ2How does LLM-based segmentation compare to human segmentation in terms of consistency?
- RQ3Can an LLM-enabled framework provide valid recall assessments based on semantic similarity to segmented events?
- RQ4Does the LLM-based recall scoring offer a scalable alternative to manual scoring without sacrificing validity?
Key findings
- LLMs can accurately identify event boundaries in narratives.
- Human segmentation is more consistent with LLMs than with other humans.
- An automated recall assessment framework using semantic similarity can estimate recall performance.
- LLM-driven methods provide a scalable alternative to manual scoring for studying perception, memory, and cognitive impairment.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.