[论文解读] Event Segmentation Applications in Large Language Model Enabled Automated Recall Assessments
这篇论文证明大型语言模型可以自动化事件分割与回忆评估,与人类分割模式一致,并为手动评分提供可扩展的替代方案。
Understanding how individuals perceive and recall information in their natural environments is critical to understanding potential failures in perception (e.g., sensory loss) and memory (e.g., dementia). Event segmentation, the process of identifying distinct events within dynamic environments, is central to how we perceive, encode, and recall experiences. This cognitive process not only influences moment-to-moment comprehension but also shapes event specific memory. Despite the importance of event segmentation and event memory, current research methodologies rely heavily on human judgements for assessing segmentation patterns and recall ability, which are subjective and time-consuming. A few approaches have been introduced to automate event segmentation and recall scoring, but validity with human responses and ease of implementation require further advancements. To address these concerns, we leverage Large Language Models (LLMs) to automate event segmentation and assess recall, employing chat completion and text-embedding models, respectively. We validated these models against human annotations and determined that LLMs can accurately identify event boundaries, and that human event segmentation is more consistent with LLMs than among humans themselves. Using this framework, we advanced an automated approach for recall assessments which revealed semantic similarity between segmented narrative events and participant recall can estimate recall performance. Our findings demonstrate that LLMs can effectively simulate human segmentation patterns and provide recall evaluations that are a scalable alternative to manual scoring. This research opens novel avenues for studying the intersection between perception, memory, and cognitive impairment using methodologies driven by artificial intelligence.
研究动机与目标
- 研究在人们在自然环境中如何感知和回忆事件及其对感知和记忆障碍的相关性的动机。
- 开发基于LLM的事件分割自动化流水线,并通过分段事件的语义相似性来评估回忆。
- 在人工注释方面验证基于LLM的分割以评估有效性,并比较人类与LLM之间的一致性。
- 展示通过LLM实现的自动回忆评分可以在记忆相关研究中实现认知评估的可扩展性。
提出的方法
- 使用基于对话的完成模型对叙述进行自动事件分割。
- 使用文本嵌入模型量化对分段事件的回忆相似性。
- 将自动分割结果与人工注释进行比对以评估有效性。
- 比较人类分割与LLM分割的一致性,以及人-人之间的一致性。
- 提供一个框架,使分段事件与回忆估计之间的语义相似性成为回忆度量。
实验结果
研究问题
- RQ1LLMs是否能够在自然叙事中准确识别事件边界?
- RQ2在一致性方面,LLM分割与人类分割相比如何?
- RQ3一个基于LLM的框架是否能够基于分段事件的语义相似性提供有效的回忆评估?
- RQ4LLM驱动的回忆评分是否能在不牺牲有效性的前提下为手动评分提供可扩展的替代方案?
主要发现
- LLMs能够在叙事中准确识别事件边界。
- 人类分割与LLMs的一致性高于与其他人类的一致性。
- 基于语义相似性的自动回忆评估框架可以估计回忆表现。
- 基于LLM的方法为研究感知、记忆和认知障碍提供了可扩展的手动评分替代方案。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。